Made a small, quick website showing GPQA benchmark scores plotted against LLM inference cost, at https://ai-benchmark-price.glitch.me/. See how much you get for your buck:
Most benchmark data is from Epoch AI, except for those marked “not verified”, which I got from the model developer. Pricing data is from OpenRouter.
All the LLMs on this graph which are on the Pareto frontier of performance vs price were released December 2024 or later...
Made a small, quick website showing GPQA benchmark scores plotted against LLM inference cost, at https://ai-benchmark-price.glitch.me/. See how much you get for your buck:
Most benchmark data is from Epoch AI, except for those marked “not verified”, which I got from the model developer. Pricing data is from OpenRouter.
All the LLMs on this graph which are on the Pareto frontier of performance vs price were released December 2024 or later...