Back to home

Agent

30 articles tagged with this topic

AgentContext Engineering

Longer Chats, Dumber AI: Agent Bottleneck Isn't Memory—It's Your PPT

An engineer argues Agents stall because they use chat logs as working memory. When artifacts become directly readable, editable, and verifiable, Agent

16h ago2 min read
AlibabaPixelle-Video

Alibaba's Pixelle-Video: 5-Minute Demo Is Hype, the Pipeline Is the Playbook

Alibaba AIDC open-sources Pixelle-Video, a pluggable short-video pipeline. The 5-minute human time is the gimmick — the orchestration architecture is

16h ago2 min read
AgentAWS

Agents need 'lockfiles' too — AI assistants break between updates, not the model

AI assistants breaking between updates isn't a model problem—it's skills, tools, and permissions drifting. AWS and OpenAI are adding version locks.

18h ago2 min read
Andrew Ngprompt engineering

AI Projects Keep Failing — Stop Blaming LLMs, Engineering Is the Real Problem

Tech team post-mortem: not the model's fault — bad prompts, unsanitized input, unvalidated output. A systemic enterprise AI failure pattern.

1d ago2 min read
DoubaoByteDance

Doubao Upgrades from Chat Tool to 'Computer Operator' — ByteDance's Most Popular Chinese AI Now

Doubao quietly completed a major upgrade: from Q&A tool to computer operator. It can directly manipulate local files, drive browsers and Feishu, and a

1d ago2 min read
AI testingAI application evaluation

Testers Become AI Quality Inspectors: Open-Source Roadmap Exposes a Talent Gap

A 13-chapter AI testing roadmap hit GitHub, gaining thousands of stars. Real signal: "who verifies AI" is becoming a new job.

1d ago2 min read
Zhipu AIAlibaba Qwen

GLM, Qwen Catch Closed-Source Leaders on Agent Benchmarks in Just Two Months

GLM and Qwen now match top closed-source models on Agent Arena Code — a result unthinkable two months ago.

1d ago2 min read
OpenAICodex

Codex plugs into Claude Code: Big-tech AI tools team up, single-vendor era ends

On Aug 24, OpenAI turned Codex into an official Claude Code plugin. Enterprise IT shifts from "pick one Agent" to "who's in the driver's seat."

2d ago2 min read
AntVAnt Group

Ant Group ships free AI-native viz tool Sive — agents reading data still early

Ant Group's AntV launches Sive, a free AI-native visualization platform — natural language to charts and reports, Agent-ready.

2d ago2 min read
ByteDanceDoubao Work

Doubao Work Agent Watches Repos and Reports — ByteDance Battles for Office

Doubao Work Agent: connectors, scheduled checks, reusable Skills for daily busywork. Real Agent traction, or ByteDance's 30-day free funnel?

2d ago2 min read
WorkBuddyCoze

600+ Agent Tutorials Open-Sourced: Chinese AI Learning Finally Gets a Roadmap

A bigtech architect consolidated 628 engineering docs, 23 WorkBuddy, and 37 Coze pieces into a three-stage open-source path. Chinese Agent learning la

2d ago2 min read
AgentStreaming Interaction

AI Agents: From Demo to Production, Streaming Is the 'Wait' Divide

For AI agents, the divide is the wait: streaming cuts first-token latency from seconds to hundreds of milliseconds, making or breaking retention.

2d ago2 min read
PLAN-AND-ACTICML 2025

AI Agents Fail 9 of 10 Long Tasks — ICML Paper Splits Roles to Hit 57%

ICML 2025's PLAN-AND-ACT splits Agents into Planner + Executor, lifting WebArena-Lite success from 9.85% to 57.58%. The fix isn't smarter AI — it's st

2d ago2 min read
DeepSeekHarness

DeepSeek Agent Tutorial Hits Chapter 10 — China's LLMs Now Chase Developers

DeepSeek's Harness Agent has a 10-chapter Chinese tutorial. The shift: top LLM firms are moving from benchmark wars to developer ecosystem grabs.

2d ago2 min read
AI evaluationbenchmark

AI Aces the Test, Fails the Job — Industry Asks What Benchmarks Really Measure

Models ace benchmarks but stumble in production. We examine a rising debate: years of AI evaluation may have measured memory, not competence.

3d ago2 min read
WorkBuddyByteDance

User Count No Longer the Power Metric — China's ToB Arena Reshuffles by Tokens

WorkBuddy's monthly traffic lead is just surface-level. From 2026, AI rankings shift from user count to Token consumption — middle-layer SaaS dollars

3d ago2 min read
DeepSeekMultimodal

DeepSeek Adds Vision at Lowest Domestic Price—But Page-Cloning Trails K3

DeepSeek launched deepseek-v4-flash-vision-exp at ¥0.05/million input tokens—lowest domestic price. Closes Agent gap, but page-cloning lags K3.

3d ago2 min read
DoubaoByteDance

Doubao Turns Phones Into PC Remotes as China's Biggest AI App Chases Codex

Doubao's Work Task Mode: phones drive PCs by voice—spreadsheets, video, trends. China's largest AI app brings agent productivity to the masses.

3d ago2 min read
DeepSeekCordis

DeepSeek Open-Sources Its Agent OS — And That Matters More Than the Code

DeepSeek's Cordis Agent framework dissected in 9 chapters — first time a top Chinese AI firm has exposed production Agent infra at this depth.

3d ago2 min read
AWSOpenSearch

AWS kills window-switching for AI debugging — Big Tech races Agent's last mile

AI delivers answers in seconds; engineers still switch browsers to verify. AWS's MCP Apps puts dashboards in AI chat — the final puzzle for enterprise

4d ago2 min read
Zhou Pu DataText2SQL

AI SQL Isn't Rare — Safe Delivery to 40,000 Businesses Is: Zhou Pu Agent

A data firm serving 40,000 distributors ships an AI SQL Agent. Real insight: enterprise AI stalls on metrics, permissions, delivery — not models.

4d ago2 min read
dshCordis

Why Agents Always Leave Junk When Swapping Tools: dsh Dissects Cordis Runtime

dsh team dissects Cordis runtime: bind every state change to its inverse for auto-rollback. Decides if AI agents can truly deploy in enterprises.

4d ago2 min read
DeepSeekdsh

DeepSeek Agent Framework Reverse-Engineered — LLM Race Moves to App Layer

DeepSeek's dsh Agent framework just got fully reverse-engineered. China's LLM firms are pushing from models to app infrastructure—real impact for deve

4d ago2 min read
QwenTongyi Qianwen

Qwen Autonomously Writes a C Compiler — Open-Source Agents Enter the Long-Haul Race

Qwen 3.6 27B built a C99 compiler autonomously in 6 weeks. Open-source handles long-haul tasks, but hallucinations and context issues still put "auton

4d ago2 min read
NVIDIAVera Rubin

NVIDIA's Vera Rubin Redefines AI Compute — But Power Costs Are the Real Story

NVIDIA Vera Rubin and Blackwell hit perf-per-watt records for Agent workloads. But Agent-driven 4x input growth is lifting compute demand.

5d ago2 min read
DeepSeekHarness

DeepSeek Harness After One Week: Open-Source Foundation, Not Claude Code Clone

DeepSeek's Harness got 95K stars in one week. Testers warn it's not a Claude Code clone but an open-source base users must assemble themselves.

5d ago2 min read
ByteDance DoubaoMCP

ByteDance Doubao MCP Connector Live: 40% CX Lift—Enterprise IT Trouble Starts

ByteDance Doubao's MCP connector went live Aug 21: AI now calls enterprise systems; one case: ~40% faster handling. MCP becomes the Agent standard.

5d ago2 min read
AgentReinforcement Learning

Why AI Agents Master Coding but Stall in Healthcare and Law

Why AI Agents work for coding but fail in medicine and law: the root cause is data structure and feedback loops, not the model itself.

6d ago2 min read
LocalLLaMAReddit

The Hidden Cliff in Local LLMs: Reddit Benchmark Reshapes Enterprise AI Math

Reddit's cHunter789 built ctx-cliff: when local LLM context exceeds VRAM, models degrade by re-reading history—breaking enterprise Agent projects.

Aug 222 min read
AnthropicClaude Code

Anthropic: 80% of Code Merged by Claude, Security Review Still Built for Humans

Claude merges ~80% of Anthropic's code, but security review still assumes human authorship—a paradigm shift every AI-coding enterprise must address.

Aug 222 min read