Local AI models & private LLMs

Artificial intelligence for your business, without ever sending your data outside.

We implement 100% local and confidential AI models. No patient or client data is used for training or sent to third-party APIs: the model runs on your server, inside your perimeter, under your control.

The problem, stated plainly

Every prompt on a public AI
is data leaving your practice.

Your team has already started using AI. Not out of malice: with a medical report pasted in to get a summary, a defense brief uploaded to check its form, a client's balance sheet fed in to have a line item explained. Every time, sensitive data of real people crosses servers you know nothing about: where they are, who touches them, how long they stay. And the responsibility for that passage, in front of the Data Protection Authority and in front of the client, remains entirely yours.

01

GDPR penalties

Uploading health or judicial data to a third-party service without a legal basis, without adequate notice and without a signed agreement with the data processor exposes you to penalties that the Regulation calculates as a percentage of turnover. In practice, the problem almost always comes from an employee who just wanted to work faster.

Art. 9 GDPR Non-EU transfers Controller's liability
02

Breach of professional secrecy

For doctors, lawyers and accountants the constraint is not only legal: it is ethical and criminally relevant. Pasting the content of a medical record or a legal filing into a public chat means sharing with a third party information covered by secrecy. No clause in the terms of use of a consumer service covers you in that conversation with your professional association.

Art. 622 Italian Criminal Code Professional codes of ethics Client confidentiality
03

Data leaks and loss of control

Once data has left, it doesn't come back. Histories exposed by a bug, compromised accounts, conversations retained by policy, content reused for training depending on the plan subscribed: these are already documented scenarios. The point is not whether the provider is serious, but that you have given up the ability to decide.

Data breach Training on inputs Notification to the Authority

"The right question is not whether to use AI in your practice. It's where the data goes when you use it."

— And it has a precise technical answer: on your server.

The solution

The same power.
Inside your walls.

A language model doesn't need someone else's cloud to work. We install AI where you decide — on your practice's server or on a machine dedicated only to you — and connect it to your documents. The answers are generated by a machine that is yours, on data that has never crossed the internet.

OPTION 01 · ON-PREMISE

The server is on site

The machine sits physically in your practice or clinic, connected to the internal network. No data leaves the building, outbound traffic is closed at the firewall level. It's the configuration for those with contractual or ethical constraints that allow no nuance.

OPTION 02 · DEDICATED SERVER

A machine only yours

A server or VPS with a reserved GPU in a European data center: unshared hardware, encrypted storage, access only from your addresses. The isolation guarantees of on-premise, without buying hardware or managing a technical room. Maintenance and updates stay with us.

OPTION 03 · PRIVATE CLOUD

Multiple sites, a single perimeter

For groups with multiple offices or branches: a private environment with separate access per site, roles and permissions per department, centralized logs. Each user sees only what concerns them, and you see who asked what.

Compliance

GDPR and AI Act,
put into practice.

The AI Act does not ban artificial intelligence in professional practices: it asks you to know what it does, on what data and under whose responsibility. A local infrastructure makes these answers simple, because everything that needs to be documented is already under your control. We hand you the configuration and the technical documentation; the final assessment is made by your DPO or your legal advisor, with papers in hand rather than with the terms of use of an American service.

PUBLIC AI (THIRD-PARTY APIs)
data → external servers → ?
legal basis: to be built
DPA: whatever the provider hands over
record of processing: incomplete

LOCAL AI (SARABI.CLOUD)
data → your server → end
Art. 28 GDPR: signed DPA
Art. 32 GDPR: encryption + audit log
Art. 50 AI Act: traceable output
Annex III: documented use

The easiest perimeter to defend is the one you never opened.

Who we install it for

Those who handle sensitive data
have different needs.

Frequently asked questions

The answers, no detours.

What hardware requirements are needed?

It depends on the model and on how many people use it at the same time. For a practice with 5 to 15 workstations, a single professional GPU with 24-48 GB of VRAM is enough: it runs mid-range models with immediate responses. For higher volumes, very long contexts or large models, you move to multi-GPU configurations with 80 GB. If you prefer not to have hardware on site, the same thing runs on a dedicated VPS with a reserved GPU: no initial investment, same isolation guarantees. We do the sizing after seeing your real use cases, not before: it's the only way not to sell you a machine bigger than needed.

How do you guarantee that patient data doesn't leave the server?

Because there is no way out. The model runs entirely on your machine and doesn't need the internet to respond: in the standard configuration outbound traffic is blocked by the firewall and only the necessary internal connections stay open. There are no calls to external APIs, so there is no channel through which data could leave. On top of this we add named access with audit logs, disk encryption and conversation retention configurable down to immediate deletion. All of it is described in the DPA we sign under Art. 28 GDPR, so the guarantee isn't just words. If you want to check in person, during the demo we show you the machine with the network unplugged: it keeps answering.

Can we use our documents to train the model?

Yes, and almost always it's better to do it without training anything. With a private RAG your documents stay in their archive, indexed inside your infrastructure: the model consults them at the moment you ask the question and cites the source. The advantage is concrete — add a document and the AI knows it immediately, remove one and it immediately stops being consultable, and every answer is verifiable. When a deeper adaptation of the language is really needed, for example to reproduce the form of your practice's filings, we proceed with a fine-tuning also run locally: the data does not leave the perimeter and the resulting model is yours, not ours.

Is a local model less capable than ChatGPT?

On the jobs that really matter to a practice, the difference is by now marginal. Summarizing reports, preparing drafts, classifying cases, searching the internal archive, answering clients: these are tasks where the latest-generation open models, connected to the right documents, hold up in the comparison without problems. And there is an advantage a public model cannot have: yours knows your practice's documents, because you would never send those to it.

How long does installation take?

On a dedicated server we generally start within a couple of weeks: analysis of use cases, sizing, installation, connecting documents, team training. With on-premise hardware, the machine's delivery time is added. The first thing we deliver, though, isn't the installation but the infrastructure demo: first you see how it works, then you decide.

What happens if the model gets an answer wrong?

It happens, as with any AI, and it's the reason we configure the system to always cite the source it took the information from: checking costs you one click. Local AI is a support tool for the professional, not a replacement for their signature — and this, besides being reasonable, is exactly what the AI Act asks for in terms of human oversight in the most delicate uses.

The first step

Watch it work.
Then decide where to put it.

We show you the infrastructure at work on documents similar to yours, with the network unplugged. You leave the demo with the hardware sizing, the cost and the DPA to have your advisor read: three concrete things, no commitment.