Multimodal
7 articles tagged with this topic
DeepSeek Adds Vision at Lowest Domestic Price—But Page-Cloning Trails K3
DeepSeek launched deepseek-v4-flash-vision-exp at ¥0.05/million input tokens—lowest domestic price. Closes Agent gap, but page-cloning lags K3.
DeepSeek Pairs Vision with Flash Pricing—Open Source Finally Fields a Card
DeepSeek shipped deepseek-v4-flash-vision-exp Aug 21 at Flash rates—first Chinese LLM matching vision and text cost; Claude Code gets free vision.
NVIDIA Brings Federated Learning to Multimodal AI — Data Stays Local
NVIDIA FLARE now supports vision-language models, letting hospitals and factories train shared AI without moving data. Not a silver bullet.
Local AI Vision Models Read Meters: Community Tests Beat Vendor Benchmarks
Locally-deployed open-source vision models now recognize real-world scenes like meter readings. A developer posted meter photos to an AI community for
Alibaba Qwen Redefines Multimodal: The Agent-First Shift Has Begun
Alibaba Qwen's livestream took "Agent First" as its theme, redefining multimodal from understanding to doing. The first major Chinese player to align
AI Workplace Penetration Hits 97% in Three Years, Yet Over 40% of Agent Projects Will Fail
IDC projects China's AI endpoint penetration will approach 97% by 2029, but Gartner warns 40%+ of agentic AI projects will be canceled. The real divid
DeepSeek Multimodal Test: Instant OCR, Fails Color Blind Cards, Coding Wins Most
DeepSeek multimodal gray test: blazing OCR but fails color blind cards, exposing visual perception gaps. Coding wins most; image-to-HTML works but tra