← AI Kitchen talk · Part 3 of 5
Our approach
- Local AI on Faculty hardware, specifically two NVIDIA DGX Sparks.
- I built a RAG proxy. Universal, protocol-based. It sits in the middle: a server between the client and the model. The RAG is in the proxy, not inside the chat app.
- It speaks the usual chat protocols, so anything that can talk to ChatGPT or Ollama can talk to it.
- Pins: the proxy injects a short cheat sheet into the prompt to steer the answer. That is not vector search. It helps stop the wrong policy from being the only thing the model sees.
- We built our own web UI to manage the RAG databases (brains). Drag and drop files. They get indexed right away. You are not rebuilding an app. Clinics, IT help desk, any topic can have a local expert.
