A 395M-parameter local model runs a single decision in 10 milliseconds on a MacBook—no network, no cloud calls. That's DecisionTune 1.0, released this week by developer Abe238. It targets the "small stuff" AI agents do every day—which tool to call, which queue a support ticket lands in, whether a request is for a refund—all solvable on-device. What we're watching is the trend: small models are starting to poach work from cloud large models.

What this is

DecisionTune 1.0 is an understanding-only model (encoder-only—it doesn't "write" text, it only "picks" options). Feed it a state, a question, and a set of options; it returns a probability for each option—essentially a scorer. On an M5 Pro MacBook with the MLX backend, median per-decision latency is 9.6ms; pure CPU runs around 65ms. Weights are 1.58GB, runs fully offline once downloaded, with SHA-256 checksums on every file. License is Apache-2.0, commercial-ready out of the gate. It also ships with an MCP server (MCP is the standard protocol local AI assistants use to call external tools), so it drops into Claude Desktop, Cursor, and other local AI workflows.

Industry view

This is a concrete sample of the small language model (SLM) trend. Microsoft Phi, Apple AFM, and Hugging Face are all betting on this path: hand high-frequency, low-complexity decisions to local models, reserve cloud large models for the tasks that actually require "thinking." Across 2,755 test cases, DecisionTune 1.0 hits 99.85% answer consistency across PyTorch, ONNX, and MLX backends—solid stability.

But there are cooler voices. One: the probabilities it returns aren't calibrated—when the model says 70% chance of a refund request, the true probability may not be 70%, so using it for automated thresholds demands caution. Two: it currently supports English only, performance on Chinese is unknown. Three: it only "picks from the options you give it"—if the options themselves are wrong, it follows along. It's a magnifying glass, not a decision-maker.

Impact on regular people

For enterprise IT: high-frequency, low-risk tasks like support ticket triage, content pre-screening, and simple routing are candidates for swapping some cloud API calls with local small models—cost drops to a fraction, and privacy improves.

For working professionals: with MCP integration, you can hand off "small decisions" to local models inside Claude Desktop, Cursor, and other local assistants—your local AI workflow gets more complete.

For the consumer market: expect more "no-cloud" AI capabilities on phones and laptops going forward. Privacy improves, but the functional ceiling drops—don't expect it to do everything.