1 article tagged with this topic
SenseTime's vision model on ByteDance's Bagel does detection, OCR, depth, segmentation. Tops benchmarks; 80GB GPU plus license keep it research-only.