Back to home

Multimodal

7 articles tagged with this topic

DeepSeekMultimodal

DeepSeek Adds Vision at Lowest Domestic Price—But Page-Cloning Trails K3

DeepSeek launched deepseek-v4-flash-vision-exp at ¥0.05/million input tokens—lowest domestic price. Closes Agent gap, but page-cloning lags K3.

3d ago2 min read
DeepSeekClaude Code

DeepSeek Pairs Vision with Flash Pricing—Open Source Finally Fields a Card

DeepSeek shipped deepseek-v4-flash-vision-exp Aug 21 at Flash rates—first Chinese LLM matching vision and text cost; Claude Code gets free vision.

Aug 222 min read
NVIDIAFLARE

NVIDIA Brings Federated Learning to Multimodal AI — Data Stays Local

NVIDIA FLARE now supports vision-language models, letting hospitals and factories train shared AI without moving data. Not a silver bullet.

Aug 192 min read
LocalLLaMAOpen-source Vision Models

Local AI Vision Models Read Meters: Community Tests Beat Vendor Benchmarks

Locally-deployed open-source vision models now recognize real-world scenes like meter readings. A developer posted meter photos to an AI community for

Aug 152 min read
QwenAlibaba

Alibaba Qwen Redefines Multimodal: The Agent-First Shift Has Begun

Alibaba Qwen's livestream took "Agent First" as its theme, redefining multimodal from understanding to doing. The first major Chinese player to align

Aug 142 min read
IDCGartner

AI Workplace Penetration Hits 97% in Three Years, Yet Over 40% of Agent Projects Will Fail

IDC projects China's AI endpoint penetration will approach 97% by 2029, but Gartner warns 40%+ of agentic AI projects will be canceled. The real divid

Aug 102 min read
DeepSeekMultimodal

DeepSeek Multimodal Test: Instant OCR, Fails Color Blind Cards, Coding Wins Most

DeepSeek multimodal gray test: blazing OCR but fails color blind cards, exposing visual perception gaps. Coding wins most; image-to-HTML works but tra

May 12 min read