Hôtel Belhoula
Room 0801 · Floor 08 · Kaoun

Self‑hosted language model

Low‑latency inference on hardware we control, so sensitive data never leaves.

The problem

Fintech data can't go to an external API, but the products still needed fast LLM inference.

What I built

  • Deployed and served a language model on dedicated servers and a rented DGX supercomputer.
  • Designed a PII obfuscation module on top of it to anonymise sensitive data for compliance.