AI Automation, RAG & MCP
5 min read
By UnlockLive IT engineering team
Diagram of a private retrieval-augmented generation system running inside a company's own cloud network with no connection to public AI services

Teams in healthcare, finance and law want the same AI help as everyone else. They want to ask questions of policies, case files, contracts and reports, and get a sourced answer in seconds. But many cannot send those documents to a public AI service. Patient records, client files and financial data come with contracts, regulators and reputations attached.

The result is often a quiet standstill. Staff are told not to use AI tools, or they use them anyway on personal accounts. Neither is good. Private RAG offers a middle path: the useful part of AI, running on infrastructure you control.

What this assistant does

Retrieval-augmented generation (RAG) means the system searches your documents first, then has a language model answer using only the passages it found. In a private setup:

  • Documents stay in your environment. The index, embedding model and language model run in your cloud account, data center or on individual laptops.
  • Questions stay in your environment too. Prompts and answers are not sent to a public AI service.
  • Answers cite their sources and respect each user's document permissions.
  • Everything is logged in your own audit system, under your own retention rules.

Deployment options

"Private" covers a range of choices. From least to most control:

  • Hosted API under business terms. A major model provider, with a data-processing agreement, regional hosting and limited retention. Simplest, but data still leaves your network.
  • Open-source model in your cloud account. GPU instances inside your own virtual network. Data stays in your account, and you manage the servers.
  • On-premise servers. Models on hardware in your own data center. Maximum control, with the most operational work.
  • On-device. Smaller models on employee laptops, for individual work with highly sensitive files.

Many organizations mix them. Low-risk content can use a hosted model, while regulated content stays on private infrastructure.

Choosing between the options

A few questions usually settle the choice:

  • Do your contracts or regulators allow this data to be processed by a third party at all?
  • If they do, which regions and which kinds of agreement are acceptable?
  • How many people will use the system, and how often?
  • Do you have staff who can run GPU servers, or will a partner do it?
  • How much answer quality are you willing to trade for control?

Your security, privacy and legal teams should agree the answers before any build starts.

How it works, step by step

  1. Ingest documents. Connectors read from your document management system, file shares or databases inside your network. Sensitive fields can be masked before indexing where they are not needed.
  2. Split and index. Documents are split into sections and turned into embeddings, numeric representations of meaning, by a locally hosted embedding model. They are stored in an encrypted vector database with permission metadata.
  3. Retrieve for each question. The system checks who is asking and searches only documents they are allowed to read.
  4. Answer with citations. A private language model writes the answer from the retrieved passages and links each source.
  5. Log and improve. Queries and sources are logged inside your environment for audit and quality review.

Tools we use

  • Backend: Python with FastAPI, deployed in your cloud account or on your servers.
  • Vector database: pgvector in your existing Postgres, or Qdrant, both self-hosted.
  • Models: open-weight families such as Llama, Mistral or Qwen, served with tools like vLLM or Ollama, plus local embedding models.
  • Infrastructure: GPU instances on AWS, Azure or Google Cloud in your chosen region, or on-premise hardware.

Check model licences, cloud pricing and GPU availability before you plan. Licence terms differ between models, and GPU capacity varies by region.

Trade-offs vs hosted APIs

Quality

Good retrieval does much of the work in document question-answering, and current open-weight models handle it well. The largest hosted models still tend to lead on complex reasoning. Test both on your own questions before deciding.

Cost

Hosted APIs charge per use. Private models cost money for hardware or GPU time whether busy or idle. At steady, high volume, private hosting can cost less. We built a private on-device AI deployment for a regulated enterprise client. It eliminated $4.2K a month of OpenAI spend, and no customer data leaves the laptop. Your numbers will depend on your volume.

Maintenance

With a private model, you own updates, security patches, monitoring and capacity planning. Budget for that time, or for a partner to handle it.

New open-weight models appear often. Swapping one in is usually simple, but each swap must be re-tested against your evaluation set before release. Treat model changes like any other production change.

What you need to get started

  • A clear use case and the document sources it needs.
  • A decision, with security and legal, on which deployment option is acceptable.
  • Access to a cloud account or servers, including GPU capacity if needed.
  • A set of real questions with known answers for evaluation.

Typical scope and timeline

A first version is typically 2 to 4 weeks, depending on sources, infrastructure and integrations. That is an estimate. Waiting for GPU capacity or security approvals can extend it, so start those early.

How we keep answers accurate

  • An evaluation set of real questions, run before every model or prompt change.
  • Citations on every answer, so experts can verify quickly.
  • Refusal when unsure, especially for clinical, financial or legal questions.
  • Monitoring of unanswered questions, feedback and response times, all inside your environment.

Risks and how we handle them

  • Privacy and compliance: private hosting helps but is not compliance on its own. HIPAA, PIPEDA and GDPR each set requirements on access, security and retention. Review them with your own counsel. See RAG privacy by design for engineering patterns.
  • Permissions: enforced at retrieval time from your identity provider, never left to the model.
  • Wrong answers: citations, refusal rules and expert review for anything that affects a client or patient.
  • Outdated documents: scheduled re-indexing and removal of retired documents.
  • Security of the stack: encryption, network isolation, patching and access logs for every component.

When not to build this

  • Your data is not actually sensitive, or a hosted API under business terms is acceptable. That is simpler and often better.
  • Your volume is low and hardware would sit idle.
  • No one can maintain the infrastructure, and you do not want a partner to.
  • You need the strongest available reasoning more than you need data control.

How UnlockLive can help

We design and run private RAG systems, from model selection and GPU sizing to permission-aware retrieval and evaluation. See our RAG development and AI and ML development services.

Related reads: the HR policy and contract Q&A assistant, the MCP server security checklist, and self-hosted n8n vs Zapier vs Make. To discuss your data and options, book a free 30-minute call.

Frequently asked questions

What is private RAG?

Private RAG is retrieval-augmented generation where the document index, the embedding model and the language model all run on infrastructure you control, such as your own cloud account, your servers or employee laptops. Documents and questions are not sent to a public AI service.

Are private open-source models as good as hosted models?

For many document question-answering tasks, current open-weight models perform well, especially when retrieval is good. The largest hosted models still tend to be stronger at complex reasoning and long, nuanced answers. Testing both on your own evaluation questions is the only reliable way to decide.

Is private RAG cheaper than using a hosted API?

It depends on volume. Hosted APIs charge per use, so they are cheap at low volume. Private models need hardware or GPU instances that cost money whether or not they are busy, plus maintenance time. At steady, high volume, private hosting can cost less; at low volume, it usually costs more.

Does running AI on our own servers make us HIPAA or GDPR compliant?

No, not by itself. Keeping data in-house reduces the number of third parties involved, which can simplify compliance. You still need access control, encryption, audit logs, retention rules, agreements and policies. Review the requirements that apply to you with your own counsel.

Can we use a hosted model with sensitive data instead?

Sometimes. Major providers offer business terms, data-processing agreements, regional hosting and options that limit data retention. For some organizations that is acceptable; for others, contracts or regulators require data to stay on their own infrastructure. That decision belongs with your security, privacy and legal teams.

How do you host an LLM on premise?

Choose an open-weight model that fits your task and hardware, run it with an inference server on your own GPU servers or a private cloud instance, and keep the document index and embedding model in the same environment. Add access control, encryption, logging and monitoring, and plan for updates. Test answer quality on your own documents before deciding.

How we can help

Talk to an engineer about your project

Tell us what you are building. We reply within one business day with a candid view on scope, approach and effort.

Book a free strategy call

Written by the UnlockLive IT engineering team. UnlockLive IT Limited works with clients through its Toronto headquarters and delivers engineering from its Dhaka delivery centre. About us

Related articles

AI Automation, RAG & MCPn8n AI Agents That Take Actions Safely: MCP Tools and Human ApprovalAI Automation, RAG & MCPAI Build vs Buy: An Honest Framework for Choosing Off-the-Shelf or CustomAI Automation, RAG & MCPStop Retyping Forms and PDFs: AI Document Extraction into Your ERP or CRM

Contact Us

Fill out the form below and our team will get back to you shortly to assist with your inquiry.