Agent
30 articles tagged with this topic
Longer Chats, Dumber AI: Agent Bottleneck Isn't Memory—It's Your PPT
An engineer argues Agents stall because they use chat logs as working memory. When artifacts become directly readable, editable, and verifiable, Agent
Alibaba's Pixelle-Video: 5-Minute Demo Is Hype, the Pipeline Is the Playbook
Alibaba AIDC open-sources Pixelle-Video, a pluggable short-video pipeline. The 5-minute human time is the gimmick — the orchestration architecture is
Agents need 'lockfiles' too — AI assistants break between updates, not the model
AI assistants breaking between updates isn't a model problem—it's skills, tools, and permissions drifting. AWS and OpenAI are adding version locks.
AI Projects Keep Failing — Stop Blaming LLMs, Engineering Is the Real Problem
Tech team post-mortem: not the model's fault — bad prompts, unsanitized input, unvalidated output. A systemic enterprise AI failure pattern.
Doubao Upgrades from Chat Tool to 'Computer Operator' — ByteDance's Most Popular Chinese AI Now
Doubao quietly completed a major upgrade: from Q&A tool to computer operator. It can directly manipulate local files, drive browsers and Feishu, and a
Testers Become AI Quality Inspectors: Open-Source Roadmap Exposes a Talent Gap
A 13-chapter AI testing roadmap hit GitHub, gaining thousands of stars. Real signal: "who verifies AI" is becoming a new job.
GLM, Qwen Catch Closed-Source Leaders on Agent Benchmarks in Just Two Months
GLM and Qwen now match top closed-source models on Agent Arena Code — a result unthinkable two months ago.
Codex plugs into Claude Code: Big-tech AI tools team up, single-vendor era ends
On Aug 24, OpenAI turned Codex into an official Claude Code plugin. Enterprise IT shifts from "pick one Agent" to "who's in the driver's seat."
Ant Group ships free AI-native viz tool Sive — agents reading data still early
Ant Group's AntV launches Sive, a free AI-native visualization platform — natural language to charts and reports, Agent-ready.
Doubao Work Agent Watches Repos and Reports — ByteDance Battles for Office
Doubao Work Agent: connectors, scheduled checks, reusable Skills for daily busywork. Real Agent traction, or ByteDance's 30-day free funnel?
600+ Agent Tutorials Open-Sourced: Chinese AI Learning Finally Gets a Roadmap
A bigtech architect consolidated 628 engineering docs, 23 WorkBuddy, and 37 Coze pieces into a three-stage open-source path. Chinese Agent learning la
AI Agents: From Demo to Production, Streaming Is the 'Wait' Divide
For AI agents, the divide is the wait: streaming cuts first-token latency from seconds to hundreds of milliseconds, making or breaking retention.
AI Agents Fail 9 of 10 Long Tasks — ICML Paper Splits Roles to Hit 57%
ICML 2025's PLAN-AND-ACT splits Agents into Planner + Executor, lifting WebArena-Lite success from 9.85% to 57.58%. The fix isn't smarter AI — it's st
DeepSeek Agent Tutorial Hits Chapter 10 — China's LLMs Now Chase Developers
DeepSeek's Harness Agent has a 10-chapter Chinese tutorial. The shift: top LLM firms are moving from benchmark wars to developer ecosystem grabs.
AI Aces the Test, Fails the Job — Industry Asks What Benchmarks Really Measure
Models ace benchmarks but stumble in production. We examine a rising debate: years of AI evaluation may have measured memory, not competence.
User Count No Longer the Power Metric — China's ToB Arena Reshuffles by Tokens
WorkBuddy's monthly traffic lead is just surface-level. From 2026, AI rankings shift from user count to Token consumption — middle-layer SaaS dollars
DeepSeek Adds Vision at Lowest Domestic Price—But Page-Cloning Trails K3
DeepSeek launched deepseek-v4-flash-vision-exp at ¥0.05/million input tokens—lowest domestic price. Closes Agent gap, but page-cloning lags K3.
Doubao Turns Phones Into PC Remotes as China's Biggest AI App Chases Codex
Doubao's Work Task Mode: phones drive PCs by voice—spreadsheets, video, trends. China's largest AI app brings agent productivity to the masses.
DeepSeek Open-Sources Its Agent OS — And That Matters More Than the Code
DeepSeek's Cordis Agent framework dissected in 9 chapters — first time a top Chinese AI firm has exposed production Agent infra at this depth.
AWS kills window-switching for AI debugging — Big Tech races Agent's last mile
AI delivers answers in seconds; engineers still switch browsers to verify. AWS's MCP Apps puts dashboards in AI chat — the final puzzle for enterprise
AI SQL Isn't Rare — Safe Delivery to 40,000 Businesses Is: Zhou Pu Agent
A data firm serving 40,000 distributors ships an AI SQL Agent. Real insight: enterprise AI stalls on metrics, permissions, delivery — not models.
Why Agents Always Leave Junk When Swapping Tools: dsh Dissects Cordis Runtime
dsh team dissects Cordis runtime: bind every state change to its inverse for auto-rollback. Decides if AI agents can truly deploy in enterprises.
DeepSeek Agent Framework Reverse-Engineered — LLM Race Moves to App Layer
DeepSeek's dsh Agent framework just got fully reverse-engineered. China's LLM firms are pushing from models to app infrastructure—real impact for deve
Qwen Autonomously Writes a C Compiler — Open-Source Agents Enter the Long-Haul Race
Qwen 3.6 27B built a C99 compiler autonomously in 6 weeks. Open-source handles long-haul tasks, but hallucinations and context issues still put "auton
NVIDIA's Vera Rubin Redefines AI Compute — But Power Costs Are the Real Story
NVIDIA Vera Rubin and Blackwell hit perf-per-watt records for Agent workloads. But Agent-driven 4x input growth is lifting compute demand.
DeepSeek Harness After One Week: Open-Source Foundation, Not Claude Code Clone
DeepSeek's Harness got 95K stars in one week. Testers warn it's not a Claude Code clone but an open-source base users must assemble themselves.
ByteDance Doubao MCP Connector Live: 40% CX Lift—Enterprise IT Trouble Starts
ByteDance Doubao's MCP connector went live Aug 21: AI now calls enterprise systems; one case: ~40% faster handling. MCP becomes the Agent standard.
Why AI Agents Master Coding but Stall in Healthcare and Law
Why AI Agents work for coding but fail in medicine and law: the root cause is data structure and feedback loops, not the model itself.
The Hidden Cliff in Local LLMs: Reddit Benchmark Reshapes Enterprise AI Math
Reddit's cHunter789 built ctx-cliff: when local LLM context exceeds VRAM, models degrade by re-reading history—breaking enterprise Agent projects.
Anthropic: 80% of Code Merged by Claude, Security Review Still Built for Humans
Claude merges ~80% of Anthropic's code, but security review still assumes human authorship—a paradigm shift every AI-coding enterprise must address.