Research · August 27, 2026

A better system around a small model beat a bigger model

I expected one thing when I started this experiment: I thought that building a better system around a small language model could, at least in some cases, make it perform better than simply moving to a larger model.

So we tested it. We ran 60,000 controlled evaluations on 600 HotpotQA questions, comparing LFM2-350M and LFM2-700M across five RAG architectures.

And in the setups we tested, that's exactly what happened. LFM2-350M consistently outperformed LFM2-700M on Token F1. With Naive RAG: 350M → 11.07% Token F1, 700M → 7.27% Token F1.

But something else happened that I didn't expect. Making the RAG pipeline more sophisticated didn't necessarily improve the result. In our Advanced RAG setup, the 350M model reached 10.17% Token F1, while average latency increased from 0.581s to 1.460s per query.

The result that surprised me most came from our Oracle RAG setup. Even when the correct evidence was provided to the model, the 350M model reached only 6.22% Token F1 in our evaluation.

That made me stop thinking about this as only a retrieval problem. In our experiments, getting the right information into the context wasn't always enough. What happens after the information reaches the model? How well can a small model actually use the context it is given?

I'm still exploring that question. The paper describing these experiments is currently under review. The implementation and experiments are available here: saiharish587/RAG-Research.

Written by Thanniru Sai Teja. More writing.