← TKF talk · Slide 10 of 18
The research
Months of failure. Then I found the papers. The literature said the same thing.
I was not the first to fail this way. The pattern I had spent months living through was already measured, named, and published. Three findings, in plain English:
- MICROSOFT EMNLP 2024 - 87.5% vs 50.4%
- RAG vs fine-tuning, head to head. Same questions, same model. RAG
- read the docs and got 87.5% right. Fine-tuning tried to remember
- and got 50.4%.
- ALLEN-ZHU & LI, ICLR 2025 - 100-1,000ร
- Exposures per fact. The number of times the model needs to see a
- fact before it can recall it at inference. I was giving it about
- No amount of clever training fixes that gap.
- PHI-4 MINI MODEL CARD - RAG > FT
- The model's own creators say so. Microsoft's documentation
- explicitly recommends retrieval over fine-tuning for factual
- recall. I had been using their model the way they said not to.
๐ค๏ธ The full path took a few more turns. Vanilla off-the-shelf RAG (worked, but messy chunks) โ custom RAG engine in Perl โ wrap it in a proxy so any tool gets it. The next slides walk through what landed.
CHIP STRIP:
- ๐ Fine-Tune fabricated ยท ๐งช RAFT needed RAG ยท ๐ Vanilla RAG messy ยท
- ๐ง Custom RAG accurate ยท ๐๏ธ RAG Proxy universal
PROOF POPUPS (in nav bar): ๐ EMNLP ยท ๐ข ICLR ยท ๐ Phi-4