One 4090, ~17GB VRAM, ten lines of commands, ten minutes to run—this week Ollama turned "local function calling" (the ability for models to trigger external APIs via structured JSON) into an out-of-the-box tool. We've noticed that 27B-tier (27 billion parameter) open-source models can now stably handle enterprise internal automation tasks on a single consumer-grade GPU. The "local Agent" (AI that autonomously invokes tools to complete multi-step tasks) has shifted from a geek toy to a viable option for enterprise IT.
What This Is
Function calling lets large models trigger external tools in a structured format—for example, "check Shanghai weather" automatically becomes a weather API call, with the result fed back into the model to generate a response. Running this locally used to require engineers to hand-roll inference frameworks and configure multi-GPU setups. Ollama's update collapses "download the model, start the service, connect tools" into a handful of commands. The Qwen3 series 27B, after Q4 quantization (a technique that compresses model size with minor accuracy loss), weighs about 17GB, fitting comfortably on a single 24GB 4090/3090. Developer benchmarks show this configuration reliably outputs compliant JSON on Chinese tool-calling evaluations, covering scenarios like ticket classification, log pre-screening, and form filling.
Industry View
Supporters argue that local deployment means data never leaves the premises, inference cost approaches zero, and long-tail tasks finally pencil out—what previously required cloud flagship models costing fractions of a cent per call for small-to-medium automation can now be covered for an entire year by one GPU's electricity bill.
But opposition is equally sharp. First, 27B still lags significantly behind GPT-4/Claude-class models in multi-step reasoning, long context, and complex tool orchestration. The gap between "can run" and "production-ready" is an engineering long tail—error rates, stability, and observability all need to be addressed. Second, security boundaries remain undefined: once a local Agent connects to internal company databases, access control, call auditing, and prompt injection (attackers hijacking model behavior through carefully crafted inputs) defenses are all blank. Third, the rapid evolution of the open-source ecosystem itself is a risk—between Qwen3.8 and the next generation, model weights and interfaces may change, making today's investments require rework in three months.
Impact on Regular People
- For enterprise IT: The "unit of accounting" for internal automation (tickets, logs, forms) shifts from "fractions of a cent per call" to "a few hundred kilowatt-hours per machine per year"—the ROI logic needs rewriting. But if compliance and auditing don't keep up, rushing into production adds new risk rather than removing it.
- For working professionals: Developers, ops engineers, and junior analysts who master Prompt writing and tool orchestration become more valuable. Pure "glue code that calls model APIs" roles will compress.
- For consumer markets: On-device AI (running directly on the device, not depending on the cloud) experiences will improve rapidly—smart speakers, vehicle infotainment, and office software can locally handle more complex tasks without connecting to the network each time. But consumer perception will lag in the short term; the shift happens first in enterprise backends.