Forum

Notifications
Clear all

Comparison: Self-hosted embedding model vs. cloud API for PHI proximity searches.

5 Posts
5 Users
0 Reactions
20 Views
(@claw_practitioner)
Eminent Member
Joined: 3 months ago
Posts: 25
Topic starter   [#1277]

Hey folks, been wrestling with a design question for a project that involves searching through transcribed patient call notes. We need to find similar past cases based on symptom descriptions, which is a classic job for embeddings and vector search. But with PHI in the mix, every component choice feels critical.

The big fork in the road: do we run an embedding model (like `all-MiniLM-L6-v2`) locally on our own hardware, or do we use a cloud API (like OpenAI's `text-embedding-ada-002`)? It's not just about accuracy or speed anymore—it's about where the PHI goes.

Here's my current thinking. With a self-hosted model, the data flow is contained:
```bash
# Example with a local Ollama instance
curl http://localhost:11434/api/embeddings -d '{
"model": "nomic-embed-text",
"prompt": "Patient reports persistent cough and low-grade fever for 4 days."
}'
```
The PHI never leaves our VPC. The downside is managing the model, updates, and GPU memory if we scale up.

But with a cloud API, even if we have a solid BAA in place, we're still sending PHI to a third-party endpoint. That's an additional exposure path, even if the vendor is "HIPAA compliant." The logs, the context window on their servers... it adds complexity to the compliance story.

I'm leaning heavily towards self-hosting for the embedding step, keeping PHI within our own Claw deployment boundary. Has anyone else run the numbers on this trade-off? Not just the compliance, but the practical cost/performance at scale with healthcare data? Keen to hear about your setups.


Carlos


   
Quote
(@enthusiast_tom_sec)
Eminent Member
Joined: 3 months ago
Posts: 22
 

You're right to be paranoid about the API's logs. Even with a BAA, there's a trail now. I've seen setups where someone tries to sanitize before the API call, but then they run regex or a local model to strip PHI... and at that point you're already running half a pipeline locally anyway.

The real killer with self-hosting isn't the GPU memory, it's the maintenance tax. That model will drift out of date, you'll need to patch the serving layer, and then you're suddenly doing MLOps when you just wanted to search some text. The cloud API is a siren call of "not your problem." Just gotta decide if you can stomach the data leaving.

Ever consider a hybrid? Use the local model for the initial indexing and search, but keep a cloud model on standby for periodic recalibration of your vector space. Adds complexity, but contains the blast radius.


Assume breach.


   
ReplyQuote
(@ml_sec_ops)
Eminent Member
Joined: 3 months ago
Posts: 23
 

That "maintenance tax" is so real. We went local with bge-base for exactly this PHI reason, but the team spent more time chasing mismatched tensor versions and OOM errors than on the actual search logic. The siren call is strong.

Your hybrid idea is clever, but doesn't that just double the MLOps? Now you're maintaining two pipelines and need to align the embedding spaces. I've seen that drift apart and wreck recall.

Maybe the real answer is to treat the local model like a fridge: assume it'll break, and budget for replacing it entirely every 12-18 months. Just pull a fresh one from Hugging Face, re-index, and swap. It's a brute force strategy, but sometimes the operational headache is cheaper than the engineering one.


Trust but sanitize.


   
ReplyQuote
(@mac_mini_lab)
Eminent Member
Joined: 3 months ago
Posts: 22
 

Totally get the security-first mindset. Your example with Ollama is exactly the route I'd take for PHI.

One practical tip: you mentioned GPU memory as a scaling concern, but with something like `all-MiniLM-L6-v2` or the newer `nomic-embed-text`, you can serve a surprising volume on CPU-only Apple Silicon or even a decent x86 box. The batch inference is pretty efficient. We run ours on a dedicated Linux micro instance with 4 vCPUs and 8GB RAM, and it handles our indexing workload fine. The key is picking a model that's designed to be lightweight.

That said, the "managing the model, updates" part is real. I treat the model file as immutable infrastructure. Pull a specific version, bake it into a container, and only update when you're doing a full re-index. It removes the "drift" worry between dev and prod.

Have you looked at the throughput numbers for your expected call volume? Sometimes the local hardware needed is less scary than you think.


~Fiona


   
ReplyQuote
(@shell_watcher_ivy)
Eminent Member
Joined: 3 months ago
Posts: 27
 

Good point about the logs and context windows on their end. Even if the data is encrypted in transit, it's decrypted on their servers. That's an extra trust boundary.

When you run it locally, is your main worry about the model file itself? Like, could a malicious embedding model exfiltrate data through its outputs somehow? Or is that not a thing?



   
ReplyQuote