AI Pilot Framework
A disciplined framework for designing, running and adjudicating a single AI pilot: choosing one low-risk task, setting the success metric before you start, running a one-third/two-thirds test with full human review, then making an honest keep, fix or kill decision.
What it is
Buying a tool is not the same as adopting AI. The teams that get value from AI do not roll it out across everything at once and hope; they prove it on one task first, under conditions honest enough to trust the result. This framework is how you run that test properly.
An AI pilot is a small, time-boxed experiment on a single workflow. You pick one task that is worth improving but safe to get wrong, decide in advance what success means and how you will measure it, run the tool on real work with a human reviewing every output, and then make a clear decision: keep it, fix it, or kill it. The discipline is in the sequencing. Most pilots fail not because the tool was weak, but because success was never defined, so the team talked itself into a soft yes once effort had been sunk.
The framework uses a one-third / two-thirds rhythm. Spend the first third of your window agreeing the rules, the metric and the boundaries. Spend the remaining two-thirds doing real work and logging what happens. A 90-day pilot splits into roughly 30 days of setup and 60 days of running, but the ratio matters more than the number: scale it to whatever horizon fits your team.
The 45 minutes below is design time. The pilot itself runs for the window you set. What you produce here is a one-page pilot plan that a busy team can actually follow and a sceptical leader can actually trust.
When to use it
Use this template when:
- You have identified a specific task or tool and want to prove its value before scaling it across the team
- Leadership is asking for evidence, not enthusiasm, before backing an AI adoption decision
- You want a defensible go/no-go you can stand behind, with the criteria fixed in advance
- You are running a Deploy or consulting engagement and need to demonstrate a use case before committing a client to rollout
Don’t use this template when:
- You have not yet chosen a task to test: run the Comms Workflow Audit first to find where the time and friction actually sit
- You are assessing the whole function’s readiness across technology, skills, process and culture: use the AI Readiness Assessment Worksheet
- You have already decided to roll the tool out regardless of results: a pilot you are not willing to kill is theatre, not a test
- The task is high-stakes or hard to reverse: a pilot belongs on work where a mistake is cheap to catch and correct
Inputs needed
- One candidate task, ideally surfaced by a workflow audit rather than chosen because it is exciting
- A way to measure the current state honestly, so the pilot has a baseline to beat
- A named person who will review every AI output before it goes anywhere
- Clarity on any boundaries that need sign-off (data handling, confidentiality, named individuals) and who owns that sign-off
- A fixed start date and decision date, agreed before you begin
The template
AI Pilot Plan
Team or function: [Name] Pilot owner: [Name and role] Task under test: [The single workflow being piloted] Tool(s): [What you are testing] Start date: [Date] Decision date: [Date, fixed in advance]
Part 1: Choose the pilot
Score two or three candidate tasks on both axes. You are looking for the task that scores high on value and high on reversibility. High value alone is a trap: it is usually also high-stakes, which is the wrong place to learn.
| Candidate task | Value (how much time or pain it removes) | Reversibility (how safe it is to get wrong) | Notes |
|---|---|---|---|
| [Task A] | High / Medium / Low | High / Medium / Low | |
| [Task B] | High / Medium / Low | High / Medium / Low | |
| [Task C] | High / Medium / Low | High / Medium / Low |
Task chosen: [The one task, stated as a specific job, not a tool] Why this one: [One sentence: worth doing, safe to get wrong]
Part 2: Define success before you start
Set one primary metric. Resist the temptation to track five: a pilot with one clear measure produces a clear decision. Then set the guardrails that must not slip while you chase it.
Primary success metric (choose one):
| Metric | Current baseline (measured, not guessed) | Target at decision date |
|---|---|---|
| [e.g. median time to a review-ready first draft] | [e.g. 45 minutes] | [e.g. under 15 minutes] |
Guardrail metrics (the things that must not get worse):
| Guardrail | Threshold it must hold | How you will check it |
|---|---|---|
| Factual accuracy | [e.g. every output human-verified, zero errors published] | [Review log] |
| Brand-voice fit | [e.g. no drop against the voice checklist] | [Spot check] |
| [Other, e.g. complaints or rework] | [Threshold] | [Method] |
Part 3: Set the boundaries
Decide where the tool is not allowed to go before it touches real work, and confirm you have the right people alongside you. These are hardest to walk back after the fact.
Off-limits for this pilot: [Confidential information, personal data, anything involving named individuals, any output that goes out without review]
| Bridge | Consulted? | Date | Note |
|---|---|---|---|
| IT / security | Yes / N/A | [Access, sandbox, approved tool] | |
| Legal / Compliance | Yes / N/A | [Data, copyright, disclosure] | |
| Data protection (DPO) | Yes / N/A | [What may and may not be entered] |
Human review rule: [Name] reviews every output before it is used, for the full pilot. No exceptions during the test.
Part 4: Name the opportunity now
Decide, before you see any results, what the freed-up time or attention will go towards. Name it now, or a success will quietly be banked as a reason to cut rather than a reason to do better work.
If this pilot succeeds and frees up [time / effort], that goes towards: [The specific higher-value work: proactive content, deeper stakeholder engagement, measurement, planning]
Does the team have the skills for that higher-value work? [Yes / Partly / Needs development, cross-reference the Capability Gap Analysis]
Part 5: Run the one-third / two-thirds timeline
| Phase | Window (default 90 days) | Focus | Done when |
|---|---|---|---|
| Setup (first third) | Days 1 to 30 | Agree the rules, capture the baseline, secure access, draft the prompts and templates | Baseline recorded, boundaries signed off, everyone knows the metric |
| Run (remaining two-thirds) | Days 31 to 90 | Real work through the tool, every output reviewed, results logged weekly | The window closes on the fixed decision date |
Scale the window to your team, but hold the ratio: roughly one-third preparing, two-thirds doing.
Part 6: The decision
On the decision date, hold the pilot to the metric you set in Part 2. Do not move the goalposts because the result was close, or because effort has been spent.
| Question | Answer |
|---|---|
| Did the primary metric hit its target? | Yes / No / Close: [figure vs target] |
| Did every guardrail hold? | Yes / No: [which slipped] |
| Was the review burden sustainable at scale? | Yes / No: [note] |
Decision:
- Keep and scale: it beat the metric and the guardrails held. Name the next step and which template supports it (workflow redesign, tool register, use policy).
- Fix and re-run: promising but not there. State the single change you will make and the new decision date.
- Kill: it did not earn its place. Record why, so the learning is not lost and the same pilot is not run again by accident.
Decision rationale (held against the original metric): [Two or three sentences]
AI prompt
Base prompt
I'm designing an AI pilot for a communications team and want it rigorous enough that the result is trustworthy.
Task I want to pilot: [DESCRIBE THE SINGLE TASK]
Why this task: [WORTH DOING AND SAFE TO GET WRONG?]
Tool I intend to use: [TOOL]
Current state: [HOW THE TASK IS DONE NOW, AND ROUGHLY HOW LONG IT TAKES]
Window available: [E.G. 90 DAYS]
Please help me:
1. Pressure-test whether this is the right task to pilot, or whether it is too high-stakes to learn on
2. Recommend a single primary success metric, and a realistic target, with a baseline I should capture first
3. Suggest 2 or 3 guardrail metrics that must not get worse while I chase the primary one
4. Split my window into a one-third setup / two-thirds run plan with clear "done when" milestones
5. Draft the keep / fix / kill decision criteria I should fix in advance
Be direct. If the metric I'm implying is a vanity measure, or the task is a bad choice for a pilot, say so.
Prompt variations
Variation 1: Pre-mortem the pilot design
Here is my AI pilot design:
[PASTE THE PILOT PLAN: TASK, METRIC, BASELINE, GUARDRAILS, TIMELINE, DECISION CRITERIA]
Assume this pilot has failed, or worse, produced a misleading success. Work backwards:
1. What are the three most likely ways this design produces a false positive (looks like a win but isn't)?
2. Is the primary metric gameable? Could it improve while the real quality drops?
3. Is the task actually reversible, or have I underestimated the stakes?
4. Where is the human-review rule most likely to quietly break under time pressure?
5. What one change would most improve the reliability of the result?
Be sceptical. I would rather find the flaw now than at the decision date.
Variation 2: End-of-pilot decision analysis
My AI pilot has reached its decision date. Help me make an honest call, held to the criteria I set at the start.
The metric I set in advance: [PRIMARY METRIC, BASELINE, TARGET]
Guardrails I set: [LIST]
What actually happened: [RESULTS: PRIMARY METRIC OUTCOME, GUARDRAIL OUTCOMES, REVIEW BURDEN, ANYTHING UNEXPECTED]
Please:
1. State plainly whether the pilot met the target I set: do not soften a near-miss into a pass
2. Recommend keep, fix or kill, with your reasoning tied to the original metric
3. If "fix", identify the single most important change and what a re-run should test
4. If "keep", outline what scaling responsibly looks like and what new risks appear at scale
5. Draft a 150-word summary I can put in front of leadership
Do not move the goalposts. If it did not hit the mark I set, tell me.
Variation 3: Turn time saved into higher-value work
My AI pilot looks likely to free up [AMOUNT OF TIME] on [TASK] for a comms team of [SIZE].
I want to decide, before the results are final, what that time should go towards, so a success strengthens the function rather than inviting a headcount cut.
Please:
1. Suggest 3 higher-value uses of the freed-up time that raise the standing of the comms team
2. For each, note the skills the team would need and whether that is a gap to close
3. Draft a short narrative I can use with leadership that frames the saving as reinvestment in better work rather than a pure cost cut
4. Flag the risk of each option so I go in with eyes open
Human review checklist
- The metric was set before the pilot started: success is defined in advance, not reverse-engineered from whatever happened
- The task is low-stakes: a mistake during the pilot is cheap to catch and easy to reverse
- A real baseline was captured: there is a measured before to compare against, not an assertion after the fact
- Guardrail metrics are defined: the things that must not slip while you chase the primary metric are written down with thresholds
- Every output is reviewed by a named human: the review rule is specific and holds for the whole pilot
- Boundaries are signed off: anyone who needs to be consulted on data, risk or compliance has been, before real work begins
- The opportunity is named in advance: what the saved time goes towards is decided before the results, not after
- The timeline holds the one-third / two-thirds ratio: enough setup, then a real run, with a fixed decision date
- Keep, fix and kill are all live options: kill is a real possibility, and the criteria will not be moved to avoid it
Example output
Anonymised and shortened. Real outputs will vary based on your inputs.
AI Pilot Plan Function: Membership body, comms team of four Task under test: First-pass answers to recurring member enquiries Tool: [LLM assistant, approved by IT] Window: 90 days (30 setup / 60 run)
Why this task: High value (member enquiries eat around eight hours a week) and high reversibility (every answer is an internal draft, always reviewed, never auto-sent).
Primary metric: Median time to a review-ready first draft. Baseline 45 minutes per enquiry, target under 15 minutes.
Guardrails: Factual accuracy human-verified with zero errors published; no member personal data entered into the tool; tone holds against the voice checklist.
Boundaries: Cleared with the Head of Membership and IT. No personal data, no auto-sending.
Opportunity named: Time saved goes to proactive member-story content, not headcount.
Decision at day 90: Median draft time landed at 12 minutes; accuracy held; one recurring tone issue fixed with a prompt change in week six. Decision: Keep, extend to two further enquiry types, and feed the standing AI use policy.
This is an illustrative example. Your pilot will reflect your own task, tool and constraints.
Tips for success
Define success before you touch the tool. The single most common reason pilots produce arguments instead of decisions is that no one agreed what a win looked like. Set the metric, the baseline and the target first. Everything else is easier once that is fixed.
Pick something you can afford to get wrong. The instinct is to test AI on your most painful, highest-value task. That is usually also your highest-stakes one, which is the worst place to learn. Choose work where a bad output is caught in review and costs nothing.
Protect the review gate. A pilot where a human reviews every output is a test of the tool. A pilot where review quietly lapses under deadline pressure is a test of your luck. Name the reviewer and hold the rule.
Name the reinvestment now. If you wait until the pilot succeeds to decide what the saved time is for, someone else may decide for you. Deciding in advance that the time buys better work, not fewer people, is what turns adoption from a cost story into a value story.
Keep it small. A pilot that tries to prove AI across three workflows proves nothing cleanly. One task, one metric, one decision. Run the next pilot after this one lands.
Write the kill criteria down. A pilot you are unwilling to stop is not a test. Deciding in advance what would make you walk away is what makes a yes worth trusting.
Common pitfalls
Vanity metrics. “The team liked it” and “it felt faster” are not results. If the metric cannot be measured against a baseline, it cannot settle the decision, and the loudest voice wins instead of the evidence.
Moving the goalposts. The result comes in just under target, effort has been spent, and the target quietly shifts to make it a pass. This is the failure that discredits every future pilot. Hold the line you set.
Skipping the baseline. Without a measured before, a pilot can only ever produce an anecdote. Capture the current state during setup, even roughly, or the whole exercise loses its evidence.
Starting with high-stakes work. Piloting AI on live crisis statements or regulated disclosures puts the learning and the risk in the same place. Learn on reversible work first, then earn your way up.
Banking a saving as a cut. A pilot that frees time without a named plan for that time invites the conclusion that the team can simply be smaller. Decide what the time is for before the numbers come in.
The pilot that never ends. With no fixed decision date, a pilot drifts into an informal, ungoverned rollout: all of the risk of adoption, none of the discipline. Set the decision date at the start and keep it.
Related templates
Need this implemented in your organisation?
Faur helps communications teams build frameworks, train teams, and embed consistent practices across channels.
Get in touch ↗