Domain
Task
All TasksImage Classification (1,593)3D Object Detection (743)Text Embedding (472)Vision-Language (285)Code Generation (125)Object Detection (106)Speech Recognition (67)Video Understanding (49)Text Generation (48)Segmentation (34)UI Understanding (14)Speech (13)Pose Estimation (11)Super Resolution (10)driver assistance (9)Font Identification (7)image editing (6)Depth Estimation (6)audio generation (6)image generation (4)robotics (3)image to text (2)video object tracking (2)gaze estimation (1)video generation (1)voice activity detection (1)
Showing 14 models
| # | Model | Domain | Task | Params ↑ | GFLOPs | VRAM | Context | Speed | Arch | Source |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ShowUI ui-understandingelement-groundinggui-agent | Multimodal | UI Understanding | 2.0B | — | 4.7 GB | 13 | — | Qwen2-VL-2B + UI-guided token selection | Curated |
| 2 | UI-TARS-2B (ByteDance) ui-understandinggui-agentelement-grounding | Multimodal | UI Understanding | 2.0B | — | 4.7 GB | 38 | — | Qwen2-VL-2B fine-tuned | Curated |
| 3 | ScreenAI (Google) ui-understandingscreenshot-qascreen2words | Multimodal | UI Understanding | 5.0B | — | 11.7 GB | — | — | PaLI-based VLM | Curated |
| 4 | Ferret-UI (Apple) ui-understandingscreenshot-analysiselement-grounding | Multimodal | UI Understanding | 7.0B | — | 16.4 GB | 63 | — | Ferret + any-resolution | Curated |
| 5 | Ferret-UI 2 (Apple) ui-understandingscreenshot-analysiselement-grounding | Multimodal | UI Understanding | 7.0B | — | 16.4 GB | — | — | Ferret-UI + multi-platform | Curated |
| 6 | UI-TARS-7B (ByteDance) ui-understandinggui-agentelement-grounding | Multimodal | UI Understanding | 7.0B | — | 16.4 GB | 75 | — | Qwen2-VL-7B fine-tuned | Curated |
| 7 | UGround (UIX-Qwen2-7B) ui-understandingelement-grounding | Multimodal | UI Understanding | 7.0B | — | 16.4 GB | 25 | — | Qwen2-VL-7B fine-tuned | Curated |
| 8 | Aguvis-7B ui-understandinggui-agentelement-grounding | Multimodal | UI Understanding | 7.0B | — | 16.4 GB | — | — | Qwen2-VL-7B fine-tuned | Curated |
| 9 | OS-Atlas-Pro-7B ui-understandinggui-agentelement-grounding | Multimodal | UI Understanding | 7.0B | — | 16.4 GB | 50 | — | InternVL2-based | Curated |
| 10 | SeeClick ui-understandingelement-groundinggui-agent | Multimodal | UI Understanding | 9.6B | — | 22.5 GB | 0 | — | Qwen-VL + GUI grounding | Curated |
| 11 | CogAgent ui-understandinggui-agentscreenshot-analysis | Multimodal | UI Understanding | 18.0B | — | 42.2 GB | — | — | CogVLM + high-res cross-module | Curated |
| 12 | UI-TARS-72B (ByteDance) ui-understandinggui-agentelement-grounding | Multimodal | UI Understanding | 72.0B | — | 168.8 GB | 88 | — | Qwen2-VL-72B fine-tuned | Curated |
| 13 | OmniParser (Microsoft) ui-understandingscreenshot-parsingelement-detection | Multimodal | UI Understanding | — | — | — | 100 | — | Detection + OCR + Icon pipeline | Curated |
| 14 | OmniParser V2 (Microsoft) ui-understandingscreenshot-parsingelement-detection | Multimodal | UI Understanding | — | — | — | — | — | Detection + OCR + Icon pipeline v2 | Curated |