Most conversations about adopting AI start with the model, the licence or the latest feature announcement. I think that is the wrong place to begin. The harder and more useful question is much simpler: what work are we trying to improve?
AI can produce an impressive demonstration very quickly. Turning that demonstration into something dependable, secure and genuinely useful is a different piece of work. That requires a clear problem, an owner, suitable data, sensible controls and an honest way of measuring whether anything has improved.
Start with the work, not the technology
A good starting point is to look for friction in the way work is currently carried out. Where are people repeating the same decision? Where is time lost moving information between systems? What is still being done manually because nobody has had the time to improve it? Where is useful information available but difficult to find?
These questions are more valuable than asking where AI can be added. They lead towards a real operational problem rather than a technology looking for somewhere to live.
I would broadly group the opportunities into three areas:
- helping people complete existing work more effectively;
- improving the speed, consistency or cost of an operational process;
- creating a service or capability that was not practical before.
The first two are normally the safest places to learn. Drafting a document, summarising an incident, searching internal knowledge or suggesting a code change can create useful experience without immediately giving an AI system authority over a live service.
Every proposed use should still have an outcome that can be measured. Time saved may be useful, but so are accuracy, rework, user satisfaction, response time and the number of occasions where a person has to correct the result. If success cannot be described before the pilot starts, it will be difficult to prove afterwards.
A good demonstration is not a production service
This is where many AI projects become stuck. A prototype works well with a carefully selected prompt, clean sample data and the person who built it standing nearby. Production introduces everything the demonstration avoided: real identities, inconsistent data, permissions, audit requirements, support ownership, cost limits and users who will try things the designer never expected.
That does not make prototyping a mistake. It means the prototype has answered only one question: can the idea work? It has not yet shown that the organisation can operate it safely and repeatedly.
Before moving further, I would want clear answers to some ordinary service-management questions. Who owns it? What information can it access? What happens when it is wrong? How is usage logged? Who reviews a security concern? Can it be stopped without disrupting the wider service? What is the support route, and what will it cost when usage grows?
The UK National Cyber Security Centre treats secure AI as a complete lifecycle covering design, development, deployment, operation and maintenance. That is a useful reminder that security is not a final check performed after the interesting development work has finished.
Decide what the AI may do
It helps to divide work into three categories. The first is repetitive, well-defined and low risk. This is where AI can often complete the task with suitable checks. The second combines automation with human judgement, where AI prepares, searches, summarises or recommends but a person remains responsible for the decision. The third relies on accountability, sensitive context or a human relationship and should remain with a person.
For most organisations, the second category is likely to be the largest. AI is valuable because it can increase what a capable person can handle, not because every human decision needs to be removed.
Human review must also be meaningful. A person clicking approve on every recommendation is not exercising judgement. They need enough context, authority and time to challenge the output. The Information Commissioner's Office makes this distinction particularly important where personal data and significant decisions are involved.
Trust comes from controls, not confidence
An AI system can sound certain while being wrong. Confidence in the wording is not evidence that the answer is reliable. Trust has to come from how the service is designed and operated.
For an assistant that only drafts text, the main control may be a person checking the result before it is used. An agent that can read company systems, change records or trigger actions needs much stronger controls. I would expect it to have its own identity, the minimum permissions required, a clear record of every action and a reliable way to suspend or revoke its access.
Read-only access should be the starting point. Write access should be introduced for specific actions only after the organisation has observed the system, tested failure cases and agreed where human approval is required. Microsoft's guidance for AI workloads similarly identifies tool access as a high-risk layer and recommends strict authentication, least privilege, audit trails and human approval for high-risk operations.
Data needs the same attention. An AI service cannot compensate for information that is inaccurate, duplicated, out of date or available to the wrong people. It may simply turn an existing data-quality problem into a faster and more convincing answer. The organisation should understand where prompts, responses and logs are stored, how long they are retained, which suppliers are involved and whether the information may be used to improve a provider's models.
The model is only one component
Model choice matters, but it is unlikely to be the lasting advantage. Models and prices change quickly. A service designed around a clear business process, controlled access and measurable outcomes can change its underlying model when there is a good reason to do so.
This is why I would avoid building an operating model around a single vendor's current lead. Different work may justify different models. A smaller and cheaper model may be sufficient for classification or summarisation, while a more capable model may be justified for a difficult reasoning task. The useful discipline is to choose the least expensive option that meets the required quality and risk level, then keep measuring it.
Cost monitoring needs to be part of the service from the beginning. Licences are visible, but consumption charges, integrations, storage, support and the time spent reviewing poor outputs can be less obvious. A pilot should establish the likely cost per useful result, not merely the cost per token or user.
Adoption is a service change
Providing access to an AI tool is not the same as achieving adoption. People need to understand what it is useful for, what information they may enter, where its limitations sit and how it applies to their particular role.
Generic training can explain the buttons. Role-specific examples show someone how the tool helps with the work they actually perform. Leaders should use the agreed tools themselves, make room for people to learn and be open about cases where AI is not the right answer.
There also needs to be a feedback route. Users will find useful patterns and unexpected failures that a project team did not predict. Capturing that information is part of operating the service, alongside monitoring quality, security, cost and changes made by the supplier.
A practical first step
I would start with one low-risk task that is frequent enough to measure. Record how it works today, including time, quality and common points of failure. Give the pilot a named owner, limit the data and permissions available to it, and decide in advance what would cause the trial to stop.
Run it for long enough to move beyond the novelty. Measure the useful outcomes, the corrections people make, the failures, the operating cost and whether staff continue to use it when they are no longer being reminded. At the end, decide whether to expand it, change it or stop it.
AI succeeds when it becomes a well-run part of the service rather than an impressive feature on the side. The model will continue to change. The need for ownership, security, good data, meaningful human judgement and measurable value will not.
