smolbenchmark

smolbenchmark · edge inference

Models that fit in 8 GB

Ranked by decode speed, tokens per joule, and heat — on your own hardware: tablets, phones, Macs, Jetsons, and Pis. Jetson is live; Pi, phones, and Mac minis are still cooking. The chart is decode tok/s vs output tok/J: fast isn’t enough if every token costs too much energy.

0 families 0 configs 0 live

Detailed reports

Full analysis and reports are available on the blogs:

Support me on Ko-fi