agentclaw

Articles

The Interview Questions That Separate a Fractional AI Officer From an AI Advisor

Lucas Brown and Sophie Adams · Aug 24, 2026 · 23 min read

Cover card reading: operator or advisor, eight questions tell you which one is across the table.

TL;DR

  • The title spread faster than the authority behind it. IBM found 76% of organizations now have a Chief AI Officer, up from 26% a year earlier, while only 61% of those officers control the AI budget.
  • One question does most of the work: what did you ship in the last twelve months that is still running, and who runs it now. A named system with a named owner is the answer. Anything else is a case study.
  • Evaluation is where the market thins out. LangChain surveyed 1,340 practitioners and 29.5% ran no evaluation of any kind on their agents, so "we spot-check the outputs" is the common answer rather than the unlucky one.
  • Gartner reckons only about 130 of the thousands of vendors claiming agentic AI are real. The same rebranding happened to people, and one call is usually the only screen you get.
  • An advisor is not a fraud. Buy one on purpose, at advisory prices, when your engineers already have capacity next quarter and somebody in the building can already sign.

Two jobs are hiding under one title, and the invoices look close enough that most buyers never find out which one they bought until month four. One person has owned something that ran without them in the room and can tell you what it cost last Tuesday. The other has read everything ever written about owning something that runs. Both interview well. Both use the same words. The questions below are the ones that come apart in their hands, and none of them takes more than a minute to ask.

Operator or advisor: what the words mean here

An operator has been accountable for a system that kept working after they stopped looking at it. An advisor has been accountable for the recommendation that such a system should exist. That is the whole distinction, and it has nothing to do with seniority. Sector experience does not settle it either. Neither does how well somebody talks about model selection, which is the thing most of these calls end up being about.

Fluency stopped being a filter about two years ago. Everyone at this level can hold a confident forty minutes on retrieval, agents, guardrails and where the market is going, because that material is free and the vocabulary moves faster than the experience does. What has not become free is the memory of a specific thing breaking at a specific time and costing a specific amount to fix.

The supply side explains why this matters now. IBM's 2026 CEO study, run with Oxford Economics across 2,000 chief executives in 33 countries, found 76% of organizations have a Chief AI Officer, up from 26% a year earlier. A seat that barely existed in 2024 is now near-universal, and the people filling it came from somewhere. Some ran things. Some rebranded.

Gartner named the vendor version of this in 2025 and called it agent washing: rebranding assistants, RPA and chatbots as agentic AI without the substance. Its estimate was that only about 130 of the thousands of vendors making agentic claims were the real thing. Nobody has run the equivalent count on individuals, and nobody will, because there is no register of who has actually shipped. There is only the interview.

Which is the awkward part of this whole page. More or less everything published about hiring for this seat was written by somebody who would be sitting on the other side of the table, so the answers that would get the author shown the door do not tend to make the list. That is the gap the rest of this fills.

One honest caveat before the questions. An advisor is not a fraud. Plenty of them are excellent, and some companies should buy exactly that: a sharp outside read, a sequenced list, an argument to take to the board. The failure is not hiring an advisor. The failure is paying operator prices, waiting for delivery that was never in the scope, and finding out in month four.

Why the interview is the only screen you get

A full-time executive search runs six weeks, four panels and a back-channel reference network. A fractional hire runs one call, maybe two, and a proposal. The compression is the point, it is why the model is cheap and fast, and it means the interview carries weight it was never designed to carry.

Reference calls help less than you would hope, because a candidate supplies references from the engagements that went well, and the engagements that went well are not where the difference between the two jobs shows up. Everybody looks like an operator on a project that worked.

The market gives you the base rate to argue against. Deloitte surveyed 3,235 business and IT leaders across 24 countries and found only 25% had moved 40% or more of their AI pilots into production, while 21% reported a mature governance model for agents and 37% were still using AI at surface level with little or no change to the underlying process. Set against that, a candidate whose history is entirely successes is describing a career that the published data says is rare. Ask them to prove it.

Three numbers to have in your head before the call

None of these are talking points. They are the base rates the answers get graded against.

of organizations now have a Chief AI Officer, against 26% a year earlierIBM Institute for Business Value, 2026 CEO Study (2026)
76%
have moved 40% or more of their AI pilots into productionDeloitte, State of AI in the Enterprise (2026)
25%
agentic AI vendors Gartner judged to be real, out of thousands making the claimGartner, agentic AI project forecast (2025)
130
IBM surveyed 2,000 CEOs across 33 countries between February and April 2026. Deloitte surveyed 3,235 leaders across 24 countries. Gartner's estimate accompanied a forecast that over 40% of agentic AI projects would be canceled by the end of 2027.

1. What is still running, and who runs it now?

Open with this one and give it room. Everything after it is either confirmation or damage control.

A good answer names things. The company or the department, the workflow, roughly when it went live, and the person who owns it now. It usually sounds boring, because production is boring. The shape of it: four hundred-odd supplier invoices a month, the AP lead runs it now, handover was in March. Boring is the tell you want. Somebody who has done this reaches for the mundane detail first because that is what they actually remember.

A bad answer describes a capability. "We built an agentic customer service platform for a logistics client." No workflow, no month, no owner, no present tense anywhere in the sentence. Listen for tense specifically. A case study lives in the past. A running system lives in the present, and somebody who owns one slips into the present without noticing.

The follow-up is one sentence: who would I call there? A yes gets you the only reference call worth making, the one where you ask what is still running rather than what the engagement was like. A no is not fatal on its own, because NDAs are real and common. It does mean you press harder everywhere else.

2. What does that system cost to run this month?

This is the cheapest disqualifier on the list, and almost nobody asks it.

The answer you want has a number in it, or names the person who watches the number, or names the line item that surprised them. Something like: about $180 a month in model calls now, though it was near $700 in week two before the retrieval step got cached. Anybody who has run a system past its first invoice has a story like that, because the first invoice is always a small shock.

"It depends on usage" is where this one dies. Which is true, and useless, and is what somebody says when they have only ever priced the build. The other bad shape is a confident build price with the running cost treated as a problem that arrives later and belongs to you.

The reason this separates cleanly is that the invoice only reaches whoever is still there when it lands. The FinOps Foundation's 2026 survey of 1,192 practitioners, covering more than $83 billion in annual cloud spend, found 98% now manage AI spend as part of the job, against 63% the year before. Cost of operation went from a specialist concern to a default one inside twelve months. Somebody selling AI leadership in 2026 who has never had to answer for a monthly number has been standing a long way from the work.

3. How do you know it is still giving the right answers?

Ask this about a specific system they named in question one, not in the abstract, or you will get a lecture on evaluation instead of an account of it.

A good answer names the eval set, how big it is, who wrote it, and at least one failure it caught before a customer did. The word regression tends to show up without prompting. So does an admission that the set was too small at first. People who run evaluation harnesses on live agents talk about them the way other people talk about smoke alarms: unglamorous, occasionally annoying, and the thing standing between them and a bad week.

A bad answer is "we spot-check the outputs." Or the answer turns into model choice, which is a different question wearing this one's clothes. Quality is not a property of the model you picked. It is a property of the checking you kept doing after you picked it.

This is the question where the market's own numbers are most useful, because the failure is so common that a weak answer is not even unusual.

Watching an agent is common. Testing it is not.

LangChain surveyed 1,340 practitioners in late 2025 on what they actually have wired up around agents in production.

Have some observability

89%

Have detailed tracing

62%

Run offline evals

52.4%

Run online evals

37.3%

Run no evaluation at all

29.5%

Fieldwork ran 18 November to 2 December 2025 with 1,340 responses. Among respondents with agents already in production the no-evaluation share drops to 22.8%, which is better and still roughly one in five. Quality was the top barrier to reaching production at 32%, ahead of latency at 20%. Cost was not the constraint.

Source: LangChain, State of Agent Engineering (2026)

4. What did you stop, and what had it cost by then?

A good answer is uncomfortable and specific. A named project, a rough spend, and the thing they were wrong about. Best case they volunteer that they were the one who had argued for it. The number matters less than the fact that a number exists, because it means somebody was tracking the spend while the decision was being made.

A bad answer is that nothing was ever stopped. Every engagement delivered. That is either a very short career or a very selective memory, and Gartner's forecast that over 40% of agentic AI projects will be canceled by the end of 2027 tells you which is more likely. In a market where a large minority of projects die, a spotless record usually means the work never got far enough to fail.

Watch for the deflection where the failure is always somebody else's: the client would not give them data, the sponsor left, the vendor underdelivered. Sometimes that is true. If it is true every time, you are hearing a habit rather than a history.

5. Who has the veto, and what happens when the answer is no?

This is the question that turns the interview around, and the response tells you more than the content of the answer.

A good answer asks you things back. Who signs. What the number is. Whether they can reject a department head's pet project without escalating to your CEO every time. Somebody who has done the job knows that the mandate is the difference between the work being possible and the work being theater, and they will want it settled before the money is.

A bad answer treats authority as your problem. No curiosity about who signs, no interest in what happens when they say no. That is somebody planning to hand over recommendations and let you carry them into the politics, which is fine if that is what you are buying and expensive if it is not.

The published numbers say this gap is real inside companies too, not just in the hiring of outsiders. IBM's study of more than 600 Chief AI Officers across 22 countries found 61% controlled the AI budget, so nearly four in ten sat in the seat without the ability to sign for the thing they owned. If you would like the full scope, cost and anti-fit picture for the seat itself, the fractional AI officer piece covers what the role owns and who should not hire one yet.

Horizontal bar chart of global Chief AI Officer figures from IBM: 80% say they get enough CEO support, 61% control the AI budget, 57% were appointed from inside the company, and 45% put building business cases first.
Support is close to universal and budget control is not, which is why the veto question is worth asking before the rate is agreed rather than after.Source: IBM Institute for Business Value, Chief AI Officer study, 2025
Show the data behind this graph
Measure, Chief AI Officers globallyShare
Say they get sufficient CEO support80%
Control their organization's AI budget61%
Were appointed internally rather than hired in57%
Name building business cases as a priority45%

6. What is on your ninety-day plan that we will not like?

You want to hear something unwelcome, and you want them to mean it. Canceling a tool you already pay for. Telling a department their pilot is not going forward this year. Spending the first three weeks on data access and permissions instead of producing anything anyone can demo at a leadership meeting.

What should worry you is a plan everybody likes. A ninety-day plan with no unpopular item is a plan that changes nothing, and Deloitte's finding that 37% of organizations are using AI at surface level with little or no process change is what that looks like at scale. Surface-level use is what happens when nobody was ever willing to be the person who said the quiet thing.

If they cannot answer without knowing more about your company, that is a reasonable objection. Give them the two facts they ask for and see whether the answer sharpens. And the plan itself has a shape you can mark against: the twelve lines it should carry by Friday of week one.

7. If we stop paying you in month four, what still works?

The answer worth having is a list of what survives. Systems your own people can operate, credentials sitting in your accounts rather than theirs, a runbook somebody on your side has actually used at least once without help. Ideally they describe a handover they have already done, including the part that went badly.

A bad answer is that everything stops, or that nobody on your side could run it. Some of that is fine in month one. As a plan it is a dependency you paid to create.

This question also flushes out where the account boundaries sit, which is worth knowing before anything is built rather than during an unpleasant month five. Ask who owns the repository, whose API keys are in use, and where the data goes.

8. Which of the things on our list should we not do at all?

Bring your actual list of AI ideas to the call and hand it over. Then ask them to remove one.

A good answer picks one, says why, and offers the cheaper fix instead: a process change, an integration between two systems you already pay for, or a person. Somebody who has run this work knows most of the value in the first year comes from three things going live and the other eleven being quietly dropped.

A bad answer is that everything on your list is worth doing and the only question is sequencing. That is the answer of somebody whose revenue is a function of the length of the list, and it is the clearest conflict of interest in this market. The roadmap-only versus build-team question goes through the same conflict from the other side, where the firm that writes the plan is also the firm that would get paid to build it.

A two-column sheet listing all eight interview questions on the left and the answer that disqualifies each one on the right, from a demo instead of a running system through to treating every idea on the buyer's list as worth doing.
Print it, or keep it open on a second screen. The right column is the part nobody being interviewed is going to volunteer.
Show the data behind this infographic
  • Ask: what did you ship in the last twelve months that is still running, and who runs it now? Disqualifying answer: a demo, a pilot, or a deck, with no named system and no named owner.
  • Ask: what does that system cost to run this month? Disqualifying answer: "it depends on usage", with no number and nobody watching the number.
  • Ask: how do you know it is still giving the right answers? Disqualifying answer: "we spot-check the outputs", with no eval set and no caught failure.
  • Ask: what did you stop, and what had it cost by the time you stopped it? Disqualifying answer: nothing was ever stopped and every engagement was a success.
  • Ask: who has the veto here, and what happens when the answer is no? Disqualifying answer: escalation to your CEO, with no willingness to own a refusal.
  • Ask: what is on your ninety-day plan that we will not like? Disqualifying answer: a plan with no unpopular item in it.
  • Ask: if we stop paying you in month four, what still works? Disqualifying answer: everything stops, or nobody on your side can run it.
  • Ask: which of the things on our list should we not do at all? Disqualifying answer: all of them are worth doing, and sequencing is the only advice offered.

How to grade the call without a scorecard

Skip the weighted rubric. You are having one conversation, not running a panel, and a spreadsheet gives a number that feels more objective than it is.

Use the first three questions as a gate instead. One should produce a name, two should produce a number, and three should produce a failure somebody caught. A candidate who clears all three has been near production. A candidate who clears none has not, whatever the rest of the call sounds like. The middle case, where two land and one does not, is where you dig, because that is usually somebody who built things and never operated them, or operated things somebody else built.

Questions four through eight are not a second gate. They tell you what kind of operator you are getting and where the friction will be. Somebody who answers four brilliantly and eight badly is honest about failure and structurally motivated to say yes to everything, which is a manageable problem if you know about it going in.

Decision flow through the first three questions. Asking what is still running branches on whether a system and owner are named; asking what it costs to run branches on whether a number or an owner of the number is given; asking how they know it is still correct branches on whether an eval set and a caught failure are named. Each no exits to a different verdict and the final yes exits to operator.
The order matters. A running-cost answer means nothing if the first question never produced a system to run.
Show the data behind this diagram
  • Question 1, what is still running and who runs it now. If no system and owner are named, the verdict is advisor: buy the advice and price it as advice.
  • Question 2, what does it cost to run this month. If no number is given and nobody is named as watching the number, the verdict is built it but never ran it: budget for the surprise.
  • Question 3, how do you know it is still correct. If no eval set and no caught failure are named, the verdict is shipped once and nobody is watching it now.
  • Clearing all three puts the candidate in operator territory, and the remaining conversation is about scope and price.

The same five subjects, answered from each side

Proof of work

An operator's answer
A named system, a named owner, a month it went live
An advisor's answer
A capability, a client logo, an outcome percentage

Running cost

An operator's answer
A monthly figure, and the line item that surprised them
An advisor's answer
A build price, and the running cost as a later problem

Quality control

An operator's answer
An eval set, its size, and a regression it caught
An advisor's answer
Spot checks, or a detour into model selection

Failure

An operator's answer
A project they stopped, roughly what it had cost, what they got wrong
An advisor's answer
No project was ever stopped, or the fault was always external

Handover

An operator's answer
Credentials in your accounts, a runbook somebody used
An advisor's answer
Continued access, and dependency framed as continuity

Neither column is a personality type. A good advisor will give operator answers on the subjects where they have operated, which is exactly why the questions are asked about one named system rather than in general.

The questions that are worth nothing

Four favorites that sound rigorous in the room and separate nobody.

Which model do you prefer. Every credible answer is "it depends on the workload" and everybody knows to say it. Model choice is also the decision most likely to be wrong by the time the contract is signed.

Do you have experience in our industry. This one is not worthless, it is just badly placed. Sector knowledge matters for regulated data, for domain vocabulary and for knowing which internal system everybody complains about. It does not tell you whether somebody can ship, and used as a first filter it removes better candidates than it keeps.

Certifications. Cloud and vendor certificates measure attendance. There is no credential anywhere that means this person has run an agent in production, which is the whole problem the interview exists to solve.

Can you show us a demo. A demo proves a happy path was built once. It says nothing about what happens at 2am on the third Tuesday, which is exactly where an operator and an advisor stop looking alike. If you want to watch something, ask to see a failure case instead, and watch what they do when it breaks in front of you.

Where we would fail our own questions

Run the same set on agentclaw, because a buyer's sheet that exempts its author is worth nothing.

Question eight is the one we would fail most easily. We build, so a list of fifteen ideas is commercially good news for us, and you should assume that pressure exists and test for it. Ask us to cut one before you ask us to build one.

Two situations should send you elsewhere. If your engineering team genuinely has free capacity next quarter, and somebody in the building can already sign the AI budget without convening a committee, you have both halves of the job already and would be paying us to sequence work you can sequence yourselves.

The other is simpler. We do not sell advisory-only engagements. Our fractional AI leadership work always arrives with people who build, because we are bad at watching a roadmap sit in a shared drive going stale. If a document is what you actually need, buy the document from somebody who is good at documents.

What this interview cannot catch

Plenty. It cannot tell you whether the person who charmed you in the interview is the person who shows up in month two, or whether they will still have your hours when a bigger client appears in month five. It cannot predict how they handle your specific politics, and politics is what kills more of these engagements than technical judgment does.

So do not let the interview be the last gate. Buy something small and time-boxed first and read the delivery rather than the pitch. Our own prices and what sits inside each tier are published for that reason: a starter build runs $1,500 to $2,500 fixed, a two-week production sprint is $5,000 fixed, and retainers start at $5,000 a month. Whoever you hire, the shape to insist on is the same. One scoped piece of work, a date, and something running at the end of it that your team can point at.

What buyers ask us about this

What questions should I ask a fractional AI officer in the interview?+

Start with what they shipped in the last twelve months that is still running and who runs it now, what that system costs to run this month, and how they know it is still giving the right answers. Those three separate an operator from an advisor faster than anything else. Then ask what they stopped, who holds the veto, what is unpopular in their ninety-day plan, what survives if you stop paying in month four, and which item on your list you should drop.

How do I tell an AI advisor from an operator?+

By whether they answer in the present tense. An operator describes systems that are running now, with owners and monthly costs attached. An advisor describes engagements that concluded, with outcomes and percentages attached. Neither is dishonest. They are different jobs, and only one of them leaves something behind that works without them.

Is it a red flag if they will not name a client?+

Not on its own. NDAs are ordinary in this work and plenty of good operators genuinely cannot name the logo. What they can always do is describe the workflow, the volume, when it went live and who runs it now, without identifying anyone. If the anonymity extends to the workflow itself, that is the red flag rather than the missing name.

Should I ask a fractional AI officer for a demo?+

It is a weak test. A demo proves a happy path was built once and says nothing about behavior under load, bad inputs or a model update. Ask to see a failure case instead, or ask what broke most recently in something they run and what the fix was. The answer to that is much harder to prepare.

How many candidates should I interview before deciding?+

Three is usually enough, and more than five stops adding information. The questions here are designed to produce a clear signal on one call rather than a ranked shortlist, so the gain from a fourth conversation is small compared with the cost of the delay. If all three fail the first gate, that is a signal about how you sourced them rather than about the market.

What if the best candidate turns out to be an advisor?+

Hire them for advice, on advisory terms, and be honest with yourself about who is going to build. That means a shorter engagement, a smaller number, and a separate answer to the delivery question before you sign rather than after. The expensive version of this mistake is paying delivery prices for a roadmap and discovering the gap when the roadmap is already three months stale.

Do I need to be technical to run this interview?+

No. Every question here is graded on whether the answer contains a specific thing: a name, a number, a date, a failure. You do not need to evaluate an architecture to notice that somebody described a workflow without ever saying who runs it. If you want a second opinion, put one of your own engineers on the call as a listener rather than an interviewer.

Run all eight on us

Bring your list of AI ideas and the questions from this page. We will answer them in order, and we will tell you which item on the list to drop before we quote anything.

Starter builds run $1,500 to $2,500, fixed. Retainers start at $5,000 a month. The audit is free either way.

Share thison Xon LinkedIn

Written by

Lucas Brown · AI Explainer Writer

I turn technical AI topics into explainers that show readers how the pieces fit together.

Playing guitar

Written by

Sophie Adams · Technical Writer

I turn complex AI concepts into step-by-step guides readers can follow as they work.

Journaling

Book audit