Let’s put AI to work where your data already lives.
I help banks, audit firms and telcos get real value from large language models without a single regulated record leaving their own building. If that is your constraint, I would genuinely like to hear about it.
Most writing about enterprise AI quietly assumes you can send your documents to somebody else’s endpoint. When you cannot, almost all of that advice stops applying. Sizing changes. Architecture changes. The security model changes completely, because the interesting attack surface moves inside your own network.
That constraint is where I work. This site is where I write it down, in public, with the code attached. If the writing has been useful to you, the work behind it is what I do for a living.
You are probably here because one of these is true
- Your data cannot leave. SAMA, GDPR, ISO 27001, or a client contract with a clause that makes the whole question moot. You need the capability anyway, and the usual playbook does not survive contact with that constraint.
- Somebody told you that you need to fine-tune. You are not sure that is true, and the people telling you have not asked what the actual failure looks like. Most of the time it is not true, and finding that out early saves an enormous amount of effort.
- There is a hardware decision in front of you. A serious GPU purchase, or a private cloud tenancy, and a sizing number you cannot quite trace. You want a second opinion before the commitment is made rather than after.
- Something already exists and nobody is sure it is sound. A pilot that worked, a vendor build that is now yours to run, or an architecture never reviewed by anyone without a stake in the answer.
Three shapes of work
In the order people usually need them. None start with a proposal. They start with a call where I try to talk you out of the expensive option.
Feasibility and sizing
A short, focused review that answers one question properly: can this run in your building, on what, and what does it cost you in capability if it does. You get a sizing model tied to your actual workload rather than a vendor’s, a hardware recommendation with the benchmark reasoning attached, a deployment topology, and a written go or no-go.
Why me: I have benchmarked and deployed on-prem inference on NVIDIA cards serving live production workloads, with the same codebase running in the cloud beside it for comparison. The numbers are measured, not quoted.
Fine-tuning or retrieval, decided
A straight answer on whether your problem genuinely needs a fine-tuned model. If it does not, you get a retrieval design instead. If it does, you get a dataset plan, an evaluation approach, and an honest account of what it will take to keep it working after launch.
Why me: I wrote the fifteen part series on fine-tuning that brought you to this site, and I still tell most clients not to do it. That is the useful part.
Architecture review and guidance
For teams already building. Architecture and model orchestration review, a vendor strategy that keeps you from being locked to one provider, security design that goes down to access control inside the vector layer, and an evaluation and rollout process your risk function will accept.
Why me: I built chunk level and prompt level access control inside a vector database for regulated documents, and a multi-model layer spanning six inference backends so no single vendor holds anyone hostage.
Four steps, no surprises
- A conversation, about thirty minutes. You describe the constraint. I ask questions and tell you honestly whether I am the right person. Sometimes the answer is that you do not need me, and I will say so.
- A written scope. What I will look at, what you will get, and when. Nothing starts until you have read it and agreed.
- The work. I stay close to your team throughout. No black box, no reveal at the end. You see the reasoning as it forms, because the reasoning is most of the value.
- A document you can hand to your board. Written for the people who have to approve the decision, not just the people who have to implement it.
Where this work has been done
Sector and outcome, without the adjectives. I do not put client names on my own website, so the engagements below are described by the shape of the problem rather than by who owned it.
Things I can name, because they are mine
These are also where I learn the things that only show up when you own the pager.
What you keep, whatever happens next
- A sizing model for your workload that your team can rerun when the workload changes.
- The reasoning behind every recommendation, written down, so it survives me leaving.
- A clear account of what was ruled out and why, which is usually the part that stops the same debate restarting in six months.
- An honest read on what this technology will and will not do for your organisation.
AWS Certified Solutions Architect Professional. Azure Data Scientist Associate. PMP. Open source at claw-code-parity and git-hook-guard. Twenty-two years across architecture, delivery and engineering leadership.
Tell me what you are up against
A few sentences is plenty. I read every one of these myself and I will come back to you personally.
If you would rather just write to me directly, I am at:
muasif80 [at] gmail [dot] com
Written out that way on purpose, to keep the scrapers off it. Click it and your mail client will open properly.