When does fine-tuning actually earn its keep?
Same conversation, three models. Type your own or pick an example, then run all three live and watch what fine-tuning buys, and where it stops mattering.
The two BART columns are self-hosted on a free autoscaling service, so they are slow on CPU and cold-start when idle. Haiku runs through the API. The Inference Cost Modeller shows where that cost difference actually bites.
Read the bottom row, not the summaries. Fine-tuning clearly lifted the base model into a usable dialogue summary, for almost no cost to run. But the frontier model matches or beats it with zero training and no hosting. At a thousand calls the cost gap is pocket change. It only becomes a decision at scale.
Generic summarization is exactly the case where a frontier model wins. So when does fine-tuning earn its keep? There are three real answers.
The decision framework →