Why pilots stall
Most legal AI pilots die of the same causes:
- No defined success criteria, so "it's promising" becomes the permanent verdict
- Too many use cases at once
- Pilot users chosen for enthusiasm, not representativeness
- No baseline, so there's nothing to compare against
- No decision date
This template forces each of those decisions before day one.
Before day 1: the pilot charter
Fill this out and get GC sign-off.
- Use case (one, maybe two): e.g., first-pass review of inbound NDAs
- Business problem: e.g., NDAs take 3 days; sales is frustrated
- Pilot users (5–10): Mix of enthusiasts and skeptics
- Executive sponsor: Name
- Pilot owner: Name (legal ops or legal engineer)
- Baseline metrics: Cycle time, touch time, volume, quality
- Success criteria: e.g., 40% faster, no drop in quality, 70% weekly usage
- Go/no-go date: Day 90
- Budget if we go: Estimated annual cost at full rollout
Days 1–30: set up and baseline
Week 1: Foundations
- ☐ Security and privacy review complete (no training on our data, retention settings confirmed)
- ☐ SSO and access configured for pilot users only
- ☐ Baseline metrics captured for the last 30–60 days
- ☐ Kickoff meeting: explain the why, the what, and the decision date
Week 2: Configuration
- ☐ Playbook or clause positions loaded
- ☐ Workflow connected to where work arrives (intake, email, CLM)
- ☐ Verification standard written: what must a human check before output is used?
Weeks 3–4: Guided use
- ☐ Short, use-case-specific training (30 minutes max)
- ☐ Office hours scheduled twice a week
- ☐ Feedback channel created (a Slack or Teams channel works)
- ☐ First usage report reviewed
Days 31–60: use it for real
Weeks 5–6: Full volume
- ☐ All pilot-scope work runs through the tool
- ☐ Track exceptions: where did AI get it wrong, and why?
- ☐ Refine playbook rules based on real misses
Weeks 7–8: Mid-pilot checkpoint
- ☐ Compare metrics to baseline
- ☐ Survey users (satisfaction, trust, time saved)
- ☐ Identify non-users and find out why
- ☐ Decide: continue as planned, adjust scope, or stop early
Days 61–90: prove it and decide
Weeks 9–10: Measure
- ☐ Final metrics versus baseline and success criteria
- ☐ Quality review on a sample of AI-assisted work
- ☐ Cost model at full rollout (seats, usage, implementation, admin time)
Weeks 11–12: Decide and plan
- ☐ Write a one-page results memo for the GC
- ☐ Go/no-go meeting held on the date set in the charter
- ☐ If go: rollout plan, training plan, and owner named
- ☐ If no-go: document what you learned and share it
Week 13: Communicate
- ☐ Share results with the department, including what didn't work
- ☐ Thank pilot users by name
Go/no-go criteria template
Go if all of these are true:
- Hit at least 80% of the primary success metric
- Quality equal to or better than baseline
- At least 60–70% of pilot users are active weekly
- Total cost is justified by time saved or spend avoided
No-go or extend if any of these are true:
- Accuracy issues that the vendor can't resolve
- Usage concentrated in one or two enthusiasts
- Security or contract terms still unresolved
The results memo (one page)
- What we tested and why
- Results versus success criteria (a simple table)
- What users said (two or three quotes)
- Cost at full rollout and expected return
- Recommendation and next steps
The bottom line
A pilot without a decision date is a subscription. Set your criteria before you start, measure against a real baseline, and make the call on day 90. Need candidates to pilot? Compare options across the legal infrastructure and contracts categories on CorporateLegal.tech.