返回首页

对比阅读

对比阅读:商汤新模型一脑干四活:识图、读字、测距、抠图 — 方向对,但离落地还差口气 与 商汤新模型一脑干四活:识图、读字、测距、抠图 — 方向对,但离落地还差口气

AEN
SenseTimeByteDanceBagel·

商汤新模型一脑干四活:识图、读字、测距、抠图 — 方向对,但离落地还差口气

What this is

This week SenseTime released SenseNova-Vision-7B-MoT: a 14-billion-parameter vision model claiming to handle object detection, OCR, depth estimation, and image segmentation all at once. It's a fine-tune of ByteDance's open-source multimodal foundation Bagel, activating 7 billion parameters per inference. The pitch is buzzy, but the hardware requirements and license terms mean it remains a research-stage product for now.

Industry view

Supporters back the "one model, multiple tasks" direction. Stacking specialized models means more deployment cost and maintenance overhead — once a unified model matures, it could cut an enterprise vision system's budget roughly in half. On benchmarks, the model does take first place in object detection and OCR, and it edges past Depth Anything V2 on depth estimation.

But skepticism is just as clear. On depth, it still loses to MoGe-2 (another depth-specialist), and on segmentation it can't beat SAM, the dedicated segmentation model. More concretely, it's only been validated on a single 80GB A800 GPU; there's no GGUF (a format for compressing large models onto consumer-grade GPUs); and mainstream open-source inference tools like llama.cpp don't yet support Bagel. The license is even more direct — it explicitly prohibits commercial use, even though the underlying Bagel foundation is actually Apache 2.0 and commercial-friendly. The top question in developer threads on Reddit: if you could actually squeeze this onto a 24GB GPU, would you use one model, or keep stacking Depth Anything, SAM, plus a small vision model?

Impact on regular people

For enterprise IT: don't put it on the procurement list in the short term. The 80GB VRAM requirement and non-commercial license mean it remains a research-community matter today.

For working professionals: just remember the trend — "unified multi-task models" are the direction of vision AI, but they're still far from everyday office work.

For consumer markets: features you use daily — phone photography, AR filters — won't change because of this model for at least a year.

BZH
商汤字节跳动Bagel·

商汤新模型一脑干四活:识图、读字、测距、抠图 — 方向对,但离落地还差口气

这是什么

商汤这周放出 SenseNova-Vision-7B-MoT:一个 140 亿参数的视觉模型,号称能干目标检测、OCR、深度估计、图像分割四件事。它基于字节跳动开源的多模态底座 Bagel 微调而来,每次推理激活 70 亿参数。想法很热闹,但硬件门槛和许可证决定了它今天还只是论文级产物。

行业怎么看

支持者看好「一个模型做多件事」的方向。多模型堆叠意味着更多部署成本、更高的维护负担,统一模型一旦成熟,能直接砍掉企业视觉系统的一半预算。论文里它确实在目标检测和 OCR 拿了第一,深度估计上也压过了 Depth Anything V2。

但反对意见同样明确:深度上仍输给 MoGe-2(另一款深度估计专用模型),分割任务上打不过专做分割的 SAM。更现实的是,它目前只在一张 80GB A800 显卡上验证过,没有 GGUF(一种把大模型压缩到消费级显卡上的格式),主流开源推理工具 llama.cpp 也还不支持 Bagel。许可证更直接——明确禁止商用,而底座 Bagel 本身其实是 Apache 2.0 可商用的。Reddit 开发者讨论里被问得最多的就是:真能压到 24GB 显卡上,你会用这一个,还是继续叠 Depth Anything、SAM 加一个小视觉模型?

对普通人的影响

对企业 IT:短期内不用排进采购清单。80GB 显门槛和非商用许可,决定了它今天仍是研究圈的事。

对个人职场:记住趋势就行——「多任务统一模型」是视觉 AI 的方向,但离日常办公还很远。

对消费市场:手机拍照、AR 滤镜这些你日常用的功能,至少一年内不会因为它有任何变化。