Start here
The problem is not too few tools, it is too many
Most firms have trialled at least one AI tool. Few have made a deliberate choice. The result is a patchwork of subscriptions, a handful of enthusiastic individuals, and a leadership team quietly unsure whether any of it is adding up.
The market is loud, fast-moving, and engineered to create urgency. Vendors promise transformation, demos look clean, and real briefs are messier. The job is to choose tools that fit your firm’s actual workflows — not to bend your workflows to fit the tools.
A wall of logos is not a strategy. A specific use case, measurable success criteria, and a governance plan are.
The map
Four categories, each with a different job
Before you evaluate any individual product, work out which category you actually need. That single decision saves the most time.
- General-purpose assistants. Capable generalists, fast to deploy across departments, but they need prompts, guardrails, and integration work before they are production-ready.
- Specialist research tools. Where general assistants summarise, research tools cite. Source traceability matters when outputs have to hold up in a client meeting or a board review.
- Enterprise platforms. Built for security, compliance, and scale first. They integrate with your cloud and identity systems and support data residency — with a corresponding cost and maintenance burden.
- Vertical solutions. Pre-loaded with domain patterns and guardrails for sectors such as legal, healthcare, and finance. You trade customisation for speed and lower risk.
The core three
General-purpose assistants worth knowing
Three general assistants cover most of what a professional services firm needs day to day.
- ChatGPT. The usual starting point. Strong on text generation, reasoning, and coding assistance, and it cuts time to first draft across content, proposals, and internal docs. Web access is immediate, the API makes it scalable, and Custom GPTs allow integration without heavy engineering. If you need a reliable generalist as a first tool, start here.
- Claude. The choice when tone consistency and document depth matter. Precise, measured writing suits client-facing work, and it holds nuance across long contracts, policies, and reports. Anthropic’s constitutional AI approach reduces risky or inconsistent outputs — a practical advantage in regulated contexts — and MCP integration lets you connect it systematically as usage scales.
- Gemini. Multimodal by default, handling text, images, and data in one flow. For Google Workspace teams it embeds assistance directly into Docs, Sheets, and Slides and reduces tool-switching. A long context window makes it capable with large documents. If your stack is Google-first, Gemini is the natural fit rather than an add-on.
Cite, do not summarise
Specialist research tools for defensible answers
When your team needs outputs it can defend, the research category earns its place.
- Perplexity. An answer agent with source links attached. Academic and news modes focus results for market scans, background briefs, and competitive Q&A, and a scheduled-task feature automates topic or sector monitoring over time.
- NotebookLM. Google’s take on the opposite approach: you upload your own documents and interrogate them, with answers cited back to your source files. Its YouTube integration processes transcripts, comments, and descriptions, and collaborative working helps teams align on analysis or onboarding content.
Security first
Enterprise platforms for regulated scale
When security and audit are the leading constraints, the platform layer is where you build.
- Azure OpenAI. OpenAI models inside the Microsoft ecosystem, with data staying in your Azure environment. Role-based access, network isolation, and private networking are configurable, which simplifies internal audits. It has recently expanded to support Claude and open-source models alongside OpenAI’s. The natural fit is any firm already running Microsoft 365, SharePoint, and Power Platform, with usage-based pricing that supports piloting by department then scaling.
- AWS Bedrock. A multi-model hub behind a single API, giving access to Anthropic, Meta, Cohere, and others without rewriting your integration as requirements change. Fine-tuning supports domain optimisation, and knowledge bases enable retrieval-augmented generation (RAG) to ground outputs in your proprietary data and reduce hallucinations. For AWS-first stacks it offers native security logging, cost controls, and model flexibility that avoids lock-in.
Narrow by design
Vertical solutions trade breadth for safety
Tools built for a specific sector carry domain guardrails already in place. You gain speed and lower output-error risk, but they cannot be treated as general assistants and integration often needs extra engineering.
- Legal. Harvey, CoCounsel, and Lexis+ AI focus on contract analysis, research, and drafting. Precision and citation behaviour are strong; data residency and integration flexibility are more limited.
- Healthcare. Platforms such as Nuance DAX support clinical documentation with safety-focused, empathic tone calibration. Governance and clinical risk owners need to be involved from the start.
- Finance. Bloomberg GPT combines domain data with numerical reasoning and real-time feeds, with compliance awareness built in. Accuracy and traceability are non-negotiable here.
- Manufacturing. Predictive maintenance, generative design, and process optimisation tools integrate with IoT data to spot issues earlier. SAP and Salesforce are both developing manufacturing-specific agent products built on ERP data.
Where ROI is won
Integration is where most rollouts stall
Choosing the right tool is step one. Making it work in your existing stack is where the return actually appears — or disappears.
The integration spectrum runs from full API control (highest flexibility, best for engineering-led teams) through to built-in connectors (faster to set up, constrained to what the tool supports natively). For operations teams without heavy development resource, platforms such as n8n, Power Automate, or Zapier provide low-code orchestration for end-to-end automation. Custom interfaces improve adoption by tailoring the experience to specific roles, and connecting your proprietary data to the AI layer is what grounds outputs at scale.
This maps directly to the KWA framework — Knowledge, Workflow, Agents:
- Knowledge. Centralise your institutional knowledge and make it AI-accessible.
- Workflow. Standardise the processes before you automate them.
- Agents. Deploy AI in production with appropriate oversight once the foundations are in place.
Skipping Knowledge and jumping straight to Agents is how firms end up with expensive pilots that never reach production.
The decision
Define these before you open a vendor conversation
A disciplined buying decision starts with four things written down.
- One specific use case and a measurable outcome. Not “improve productivity” — something concrete, such as cutting proposal first-draft time by 50% or reducing client reporting effort by eight hours per week.
- Your non-negotiable constraints. Security, data residency, existing integrations, and compliance requirements. Know these before the first sales call.
- A pricing estimate against expected ROI. Per-model pricing shifts frequently, so estimate volumes early to avoid budget surprises at scale.
- A plan to test, train, and govern. Run a live test, train users, and publish governance. A tool no one uses is not an investment — it is a cost.
Your buying checklist reduces to three criteria: security and compliance, integration with your existing stack, and total cost of ownership over a realistic deployment horizon.
Not optional
Governance is the last step to deploy and the first to plan
The OECD AI principles give a practical governance north star that governments already reference. Five translate directly into operational decisions for a professional services firm:
- Inclusive growth. AI should increase your team’s capacity, not silently reduce it.
- Human-centred and fair. Keep people in the loop to catch bias and maintain client trust.
- Transparency and explainability. Show how decisions are made; use AI to surface information, not to make opaque recommendations.
- Robust, secure, and safe. Assume AI will make mistakes and design for real-world failures.
- Accountability. Assign clear owners, log actions, and ensure a human is responsible for outputs.
This is the Govern step in the 5 Steps for AI Leadership framework — the approved tools, data handling rules, and risk register. It is the last step to deploy but the first to plan for. Firms that skip it early spend disproportionate time unwinding ad-hoc decisions when they try to scale.
The controlling idea
Fitness for purpose beats brand recognition
There is no single AI platform that fits every firm. The right choice depends on your team’s workflow, your existing stack, your compliance constraints, and the specific outcomes you are trying to achieve. The mistake most firms make is chasing brand recognition or vendor momentum rather than fitness for purpose.
Start with one workflow. Deploy it properly. Measure it. Then compound. That is how AI stops being a patchwork of subscriptions and starts being infrastructure your team can rely on.