How to use this scorecard
- Set your weights first. Before any demo, agree internally on how much each category matters. Adjust the suggested weights below to fit your priorities.
- Score each question 0 to 3. 0 = no or unacceptable, 1 = partially, 2 = meets needs, 3 = exceeds needs.
- Demand evidence. "Yes" with a document, a live demo, or a contract clause beats "yes" in a slide.
- Use your own documents. Run the pilot on your real contracts and questions, not the vendor's sample set.
- Multiply and total. Category average × weight = weighted score. Compare vendors on totals, then read the notes.
Suggested weights
- 1. Use-case fit and accuracy: 25%
- 2. Security, privacy, and data handling: 20%
- 3. Privilege and legal defensibility: 10%
- 4. Integration and workflow: 15%
- 5. Pricing and commercial terms: 15%
- 6. Adoption and support: 10%
- 7. Vendor viability and roadmap: 5%
1. Use-case fit and accuracy (25%)
- Which of our top three use cases does the product handle out of the box?
- What does accuracy look like on our documents in a pilot?
- Does the output cite or link to its sources so a lawyer can verify it?
- How does it handle our playbooks, clause positions, and fallback language?
- What happens with poor scans, foreign-language documents, or unusual formats?
- Can it take actions (agentic workflows), and can we control which actions need human approval?
- How does the product signal low confidence or uncertainty?
2. Security, privacy, and data handling (20%)
- Do you train any models on our data, inputs, or outputs? (Get it in the contract.)
- Which underlying model providers do you use, and what are their data terms?
- Where is our data stored and processed? Can we choose the region?
- What is the default retention period for prompts, outputs, and uploads, and can we change it?
- What certifications do you hold (e.g., SOC 2 Type II, ISO 27001)?
- Do you support SSO, role-based access, and ethical walls?
- How quickly do you notify us of a security incident?
- Can we get audit logs of who did what, and export them?
3. Privilege and legal defensibility (10%)
- Do your terms treat our content as confidential and restrict disclosure to third parties?
- Can we restrict use to named users acting at counsel's direction?
- Can AI conversations and outputs be preserved and collected for legal holds?
- How do you respond to subpoenas or government requests for customer data?
- How do you support EU AI Act transparency or other regulatory obligations that apply to us?
4. Integration and workflow (15%)
- Does it work inside Word, Outlook, Teams, or Slack, where our lawyers already are?
- Does it integrate with our CLM, DMS, matter management, or e-billing systems?
- Is there an API, and is it included in the price?
- Can we configure workflows without vendor professional services?
- How does data flow back into our systems of record?
5. Pricing and commercial terms (15%)
- Is pricing per seat, per credit, per document, or usage-based?
- What happens if we exceed usage limits: throttling, overage fees, or a forced upgrade?
- Can we see usage by user and workflow in real time?
- Is pricing locked for the full term, or only year one?
- If your model provider raises prices, who absorbs it?
- What's included versus extra: implementation, training, integrations, premium models?
- What are the exit terms, and how do we get our data back?
6. Adoption and support (10%)
- What does implementation look like, and how long does it take for a team our size?
- Who from your team will configure our playbooks, and will they transfer knowledge?
- What training do you provide for lawyers versus admins?
- What adoption metrics can we see after go-live?
- What does support look like after the first 90 days?
7. Vendor viability and roadmap (5%)
- How is the company funded, and how many in-house legal customers do you have?
- What's on the 12-month roadmap that matters to us?
- Can we speak with two reference customers of similar size and industry?
Red flags
- Refuses to put "no training on your data" in the contract
- Can only demo on their own sample documents
- Can't explain their pricing in one sentence
- No audit logs or usage visibility
- References are all law firms when you're an in-house team
The bottom line
A scorecard won't make the decision for you, but it will make sure you're comparing vendors on the same terms and asking the uncomfortable questions before signature. When you're ready to build your shortlist, browse the generative AI legal assistants and contract intelligence tools on CorporateLegal.tech.