← TKF talk · Slide 10 of 18

The research

The research
← PreviousNext →

Months of failure. Then I found the papers. The literature said the same thing.

I was not the first to fail this way. The pattern I had spent months living through was already measured, named, and published. Three findings, in plain English:

  1. MICROSOFT EMNLP 2024 - 87.5% vs 50.4%
  2. RAG vs fine-tuning, head to head. Same questions, same model. RAG
  3. read the docs and got 87.5% right. Fine-tuning tried to remember
  4. and got 50.4%.
  1. ALLEN-ZHU & LI, ICLR 2025 - 100-1,000ร—
  2. Exposures per fact. The number of times the model needs to see a
  3. fact before it can recall it at inference. I was giving it about
  4. No amount of clever training fixes that gap.
  1. PHI-4 MINI MODEL CARD - RAG > FT
  2. The model's own creators say so. Microsoft's documentation
  3. explicitly recommends retrieval over fine-tuning for factual
  4. recall. I had been using their model the way they said not to.

๐Ÿ›ค๏ธ The full path took a few more turns. Vanilla off-the-shelf RAG (worked, but messy chunks) โ†’ custom RAG engine in Perl โ†’ wrap it in a proxy so any tool gets it. The next slides walk through what landed.

CHIP STRIP:

PROOF POPUPS (in nav bar): ๐Ÿ“Š EMNLP ยท ๐Ÿ”ข ICLR ยท ๐Ÿ“ Phi-4