What this is
Here's how Maggie Appleton has been opening her days lately: she checks the pull requests her AI agent—an assistant that autonomously completes multi-step tasks—produced overnight, scans them, merges if nothing's broken, and calls this workflow "implementation as an overnight job." What once took a full day of typing now requires only a well-written spec handed to an agent.
But she quickly clarifies: this approach only holds up inside prototyping teams, because prototypes permit "good enough." Move it into a production environment where you're accountable for quality, and the method breaks down.
More telling are the tasks she flags that agents can't pick up: paper sketches (early ideas aren't even in language form yet—agents are nearly useless here); visual judgment ("should the border be 10% or 15% gray?" should be decided by looking at a swatch, not by getting asked the 36th question by an AI); team alignment (one person wrangling a herd of agents moves fast, but software is built by teams).
How the industry sees it
Supporters read Maggie's practice as a new workflow going mainstream—humans figure out what they want, AI does the legwork. Anthropic, OpenAI, and Cursor are all betting that agents can independently complete longer and longer tasks.
Pushback comes in two flavors.
The first from production frontlines: prototype "good enough" and product code "must be stable" are two different standards. Porting the former's workflow to the latter leaves a pile of unreviewed code, and technical debt—maintenance costs you'll pay later—snowballs.
The second is sharper—Maggie coined a term herself: capability gaslighting. The model crushes certain tasks, then flubs similar ones, with no way to predict which side you'll land on next. She admits she's sometimes missed outright model failures because "Opus, Claude's strongest version, couldn't possibly get this wrong." This is the lived experience of what Ethan Mollick calls the "jagged frontier"—what Maggie sees as the biggest hidden risk in shipping agents today.
And then there's the underestimated problem: team alignment. The title of one of her talks was literally "One Developer, Two Dozen Agents, Zero Alignment." Today's agentic tools are almost all single-player and local—nobody has an answer for how teams coordinate.
What this means for everyone
For enterprise IT: worth reassessing how we split time between "writing code" and "alignment discussion." As code generation costs collapse, the bottleneck shifts from "who can code" to "who can clearly define what to build." Specs, decision logs, and reviews may end up more valuable than the code itself.
For individual careers: "capability gaslighting" is a phrase every AI user should remember—the most impressive AI success that impressed us most may be the setup for our next failure. Maggie recommends keeping judgment calls close to the chest and leaving a decision trail every time an agent takes on work.
For consumer markets: don't expect "AI ships you a usable product end-to-end" at scale anytime soon. Prototypes can finish overnight, but between prototype and product sit judgment, alignment, and user validation—the parts AI can't pick up. That road is longer than people think.