This week, an engineering guide about GPT-6 circulated in the Juejin developer community. The most eye-catching figure is a multiplier: the inference cost between flagship Astra and value-tier Sol differs by 5x (for the same task volume). This isn't technical bravado—it's the AI industry's first explicit signal to users that not every task needs the most expensive model.
What this is
According to this guide, GPT-6 abandons the "one biggest model does everything" strategy and splits into three tiers: Astra (flagship) handles architectural design, complex code refactoring, and scientific reasoning; Sol (value workhorse) approaches Astra's capability at only 1/5 the cost, taking on most production high-concurrency tasks; Luna (volume model) targets data cleaning, simple classification, and low-latency scenarios.
Companion capabilities include: inference intensity can be dynamically switched mid-conversation (Low to Extra High) without breaking cache; Mid-turn steering lets you issue correction commands while long-running tasks execute; asynchronous tool calling lets the AI process other things in parallel while waiting on slow tasks; native multi-agent collaboration (multiple AIs dividing labor to complete a task), with sub-agents arbitrated by Astra. The entire lineup also supports Computer Use—letting the AI operate the computer screen, paired with Playwright (browser automation tool) and PyAutoGUI (desktop automation tool) for end-to-end automation.
The biggest cost-side win is prompt caching: when reusing the same prompt segment, input costs drop 95%.
Industry view
Supporters argue that three-tier pricing combined with dynamic inference adjustment finally makes long-running tasks and multi-agent workflows economically viable—where running a multi-agent task used to burn through the budget, now it covers only a Sol-tier simple call. This will significantly lower the bar for AI Agent (AI programs that autonomously complete multi-step tasks) adoption in enterprises.
But we note at least three counterpoints:
- Unpredictable costs. Tiered pricing is essentially "harvest by need"—bills jump the moment task complexity shifts, making budgeting hard for small and mid-sized teams.
- Multi-agent collaboration's backlash. When sub-agents conflict, an Astra-tier model is needed for arbitration—and arbitration costs may eat the savings from tiered pricing.
- Computer Use's security boundaries. Letting AI operate computers means "which operations auto-execute and which need human approval" must be clearly delineated—but the industry currently has no consensus standard.
Impact on regular people
For enterprise IT: The old approach of "one big model for everything" will give way to "model routing"—simple tasks to Luna, complex ones to Astra. Budget structure will shift from "per-token billing" to "tiered by task complexity."
For individual professionals: Learning "tiered usage" becomes a new skill. Daily emails and document organization don't need flagship models—image recognition and deep analysis do. Those who know how to invoke on demand will save significant time and subscription costs compared to "always use the strongest."
For the consumer market: AI product feature segmentation will accelerate. Free tiers invoke base models, professional tiers unlock advanced capabilities—same logic as cloud storage and membership subscriptions. Average consumers won't feel much short-term, but in two to three years, "which models your AI subscription can run" will become a new dimension in product selection.