Regression testing for local LLMs
Benchlocal runs the same test suite every time you install a new model, so you can see whether it is actually better for your work. Everything runs on your own machine.
What it does
- Pick a model, press Run. It talks to your local Ollama server, whether on this computer or another machine on your network (an RTX 3090 box, for example).
- Scores what matters for agents: tool calling (did it call the right function with the right arguments, or correctly call nothing?), answer matching, and generation speed (tokens per second).
- Compares runs side by side so you can see the effect of a new model, a different quantization, or thinking on/off.
- Your own tasks. Write tasks as small JSON files. They stay on your disk and never enter anyone's training data.
- Live progress and a Stop button for slow, thinking-heavy models.
Private by design
The only network destination is the Ollama URL you configure. No telemetry, no accounts, no uploads. License keys are verified offline. See Privacy.
Who it is for
People who run local models and need to decide between them with evidence: developers building local agents, and freelancers who need to justify a model choice to a client.
Status
Pre-release. Tested on macOS. Windows and Linux builds are set up but have not been tested on real machines yet. The sample tasks are written in Japanese (business and tool-calling scenarios); your own tasks can be in any language.