WritingmateWritingmate
AI Model

Bytedance: UI-TARS 72B

Bytedance research
Text Generation
Vision
About UI-TARS 72B

UI-TARS 72B is an open-source multimodal AI model designed specifically for automating browser and desktop tasks through visual interaction and control. The model is built with a specialized vision architecture enabling accurate interpretation and manipulation of on-screen visual data. It supports automation tasks within web browsers as well as desktop applications, including Microsoft Office and VS Code.

Core capabilities include intelligent screen detection, predictive action modeling, and efficient handling of repetitive interactions. UI-TARS employs supervised fine-tuning (SFT) tailored explicitly for computer control scenarios. It can be deployed locally or accessed via Hugging Face for demonstration purposes. Intended use cases encompass workflow automation, task scripting, and interactive desktop control applications.

Specifications
Provider
Bytedance research
Context Length
32,768 tokens
Input Types
text, image
Output Types
text
Category
Other
Added
3/26/2025

Use UI-TARS 72B and 200+ more models

Access all the best AI models in one platform. No API keys, no switching between apps.