Room 0801 · Floor 08 · Kaoun
Self‑hosted language model
Low‑latency inference on hardware we control, so sensitive data never leaves.
The problem
Fintech data can't go to an external API, but the products still needed fast LLM inference.
What I built
- Deployed and served a language model on dedicated servers and a rented DGX supercomputer.
- Designed a PII obfuscation module on top of it to anonymise sensitive data for compliance.