AI, Cloud & DevOps
AI features that solve a real job, on cloud infrastructure that stays up and stays affordable.
- Average cloud spend reduction
- -38%Average cloud spend reduction
- Deployment frequency
- DailyDeployment frequency
- Mean time to recovery
- <15 minMean time to recovery
How we approach this
We help teams put AI into production without betting the company on it — starting from the workflow you want to improve, not the model you want to use. Alongside that we handle the unglamorous half: infrastructure as code, observability, cost control and deployment pipelines that let you ship on a Friday.
Start from the job, not the model
The question we ask first is which task currently takes a person too long, and what a good answer to it looks like. If nobody can describe what a good answer looks like, the feature cannot be evaluated, and a feature that cannot be evaluated cannot be safely improved.
A surprising number of requests that arrive as AI projects turn out to be search problems, or data-quality problems, or a form that asks for the wrong things. We will say so when that is the case, because shipping a language model over a broken process just makes the process harder to see.
Evaluations before features
Before building the pipeline we build the scoring set — a few hundred real examples with known-good answers, drawn from your own data. Every subsequent change is measured against it, which turns prompt and retrieval work into engineering rather than guesswork.
This is the step teams skip, and it is why so many AI features feel impressive in a demo and unreliable in production. Without a scoring set you cannot tell whether a change helped, and you certainly cannot tell whether upgrading the model broke something.
Humans stay in the loop where it matters
For anything touching money, health, legal exposure or someone's application, we design the interface so the model proposes and a person decides. Citations back to source documents, confidence signals and an obvious path to disagree are part of the design, not an afterthought.
That is partly a safety position and partly a practical one: reviewers who can see why an answer was produced correct it quickly, and their corrections become the next round of evaluation data.
The infrastructure underneath it
Infrastructure as code from the first environment, so staging and production cannot drift apart and neither can be rebuilt only by the person who originally created it. Terraform for the estate, containers for the workloads, and a deployment pipeline anyone on the team can run.
Cost control is treated as a first-class requirement rather than a quarterly panic. We tag resources, set budget alerts, right-size instances against real usage and cache aggressively around model calls — the single largest lever on an AI feature's running cost.
What you get
- AI feature scoping and model evaluation
- RAG pipelines and vector search
- LLM integration, evals and guardrails
- Cloud architecture on AWS, GCP or Azure
- Infrastructure as code and CI/CD
- Monitoring, alerting and cost optimisation
Is this the right service for you?
We would rather tell you no early than take on work we would do badly. Here is where this service earns its keep, and where it doesn't.
A good fit if
- A high-volume task where people read documents to make a decision
- Support or operations workloads with repetitive triage
- A cloud bill growing faster than your usage
- Deploys that are manual, scary, or only one person can do
Probably not if
- You want AI in the product mainly for the announcement
- The underlying data is too inconsistent to retrieve from yet
- Training a foundation model from scratch — that is not our work
How a ai, cloud & devops project runs
- 01
Pick one workflow
We identify the single task with the best ratio of volume to complexity and define, in writing, what a correct answer looks like.
- 02
Build the eval set
A few hundred real examples with known-good answers. This becomes the scoreboard every later change is measured against.
- 03
Pipeline and interface
Retrieval, prompting and guardrails, wrapped in a review interface that shows citations and makes disagreeing easy.
- 04
Deploy and watch
Infrastructure as code, budget alerts, latency and quality monitoring, and a rollback path that does not require the person who built it.
Ways to work with us
Starting prices, not quotes. You get a fixed number in writing after discovery — these are here so you can tell early whether we're in your range.
AI feasibility sprint
We take one candidate workflow, build the evaluation set from your data and a rough pipeline against it, and report honestly on whether it is worth building.
- From
- $12,000
- Typical timeline
- 3 weeks
- Best for
- Deciding before committing a budget
Cloud & cost review
An audit of your infrastructure, deployment process and spend, returned as a prioritised list of fixes with the saving and effort attached to each.
- From
- $7,500
- Typical timeline
- 2 weeks
- Best for
- Bills or deploys that have got out of hand
Production AI feature
A retrieval pipeline, review interface, evaluation harness and monitoring, deployed on infrastructure defined as code.
- From
- $70,000
- Typical timeline
- 3–5 months
- Best for
- Putting a real workflow into production
AI, Cloud & DevOps questions
The ones that come up on nearly every call for this service.
- Will our data be used to train someone else's model?
- Not with the providers we default to. Anthropic and OpenAI's business API tiers both exclude API traffic from training by default, and we configure zero-retention options where they are available. If your compliance position requires it, we will architect for a model running inside your own cloud account instead, and tell you what that costs in both money and quality.
- How do you stop it making things up?
- Three ways, in order of effectiveness: ground every answer in retrieved documents rather than the model's memory, require citations so a reviewer can check the source in one click, and keep a human approving anything consequential. We also run the evaluation set on every change, so a regression in factual accuracy shows up before release rather than in a complaint.
- What does it cost to run each month?
- Model inference is usually the largest line and the most controllable. Aggressive caching, routing simple requests to smaller models and trimming context are the main levers, and together they routinely cut inference spend by more than half. We model expected monthly cost during the feasibility sprint so you see the running number before you commit to the build.
- Do you do DevOps work without the AI part?
- Regularly. Cloud architecture, Terraform, CI/CD, monitoring and cost reduction are standalone engagements. The two sit together on this page because AI features fail in production for infrastructure reasons more often than for model reasons.
- Which cloud should we be on?
- Whichever one your team can operate. We work across AWS, GCP and Azure and have no incentive to prefer any of them. If you are already somewhere, the cost of migrating almost never beats the cost of tuning what you have — we will tell you when it does.
Need ai, cloud & devops help?
Send us the brief, the half-written brief, or just the problem. We'll come back with what we'd do and what it costs.
