What hardware requirements are needed?
It depends on the model and on how many people use it at the same time. For a practice with 5 to 15 workstations, a single professional GPU with 24-48 GB of VRAM is enough: it runs mid-range models with immediate responses. For higher volumes, very long contexts or large models, you move to multi-GPU configurations with 80 GB. If you prefer not to have hardware on site, the same thing runs on a dedicated VPS with a reserved GPU: no initial investment, same isolation guarantees. We do the sizing after seeing your real use cases, not before: it's the only way not to sell you a machine bigger than needed.
How do you guarantee that patient data doesn't leave the server?
Because there is no way out. The model runs entirely on your machine and doesn't need the internet to respond: in the standard configuration outbound traffic is blocked by the firewall and only the necessary internal connections stay open. There are no calls to external APIs, so there is no channel through which data could leave. On top of this we add named access with audit logs, disk encryption and conversation retention configurable down to immediate deletion. All of it is described in the DPA we sign under Art. 28 GDPR, so the guarantee isn't just words. If you want to check in person, during the demo we show you the machine with the network unplugged: it keeps answering.
Can we use our documents to train the model?
Yes, and almost always it's better to do it without training anything. With a private RAG your documents stay in their archive, indexed inside your infrastructure: the model consults them at the moment you ask the question and cites the source. The advantage is concrete — add a document and the AI knows it immediately, remove one and it immediately stops being consultable, and every answer is verifiable. When a deeper adaptation of the language is really needed, for example to reproduce the form of your practice's filings, we proceed with a fine-tuning also run locally: the data does not leave the perimeter and the resulting model is yours, not ours.
Is a local model less capable than ChatGPT?
On the jobs that really matter to a practice, the difference is by now marginal. Summarizing reports, preparing drafts, classifying cases, searching the internal archive, answering clients: these are tasks where the latest-generation open models, connected to the right documents, hold up in the comparison without problems. And there is an advantage a public model cannot have: yours knows your practice's documents, because you would never send those to it.
How long does installation take?
On a dedicated server we generally start within a couple of weeks: analysis of use cases, sizing, installation, connecting documents, team training. With on-premise hardware, the machine's delivery time is added. The first thing we deliver, though, isn't the installation but the infrastructure demo: first you see how it works, then you decide.
What happens if the model gets an answer wrong?
It happens, as with any AI, and it's the reason we configure the system to always cite the source it took the information from: checking costs you one click. Local AI is a support tool for the professional, not a replacement for their signature — and this, besides being reasonable, is exactly what the AI Act asks for in terms of human oversight in the most delicate uses.