Dola Seed 2.0 Lite goes omni-modal and strengthens the Agent, the Coding, and the GUI.

BytePlus releases Dola Seed 2.0 Lite, an omni-modal AI model with upgraded Agent, Coding, and GUI capabilities, available on the ModelArk platform.

BytePlus is releasing an update for Dola Seed 2.0 Lite, the intermediate version of ByteDance's Seed family distributed internationally, making it the first omni-modal model in the series. The model now natively includes video, image, audio, and text within a single architecture, and accompanies this transition to omnimodality with improvements on three fronts: Agent, Coding, and GUI.

In cross-modal reasoning, the model can verify consistency between a soundtrack and a video, identify key moments from a natural language instruction, track people or events over time, and analyze emotion, ambient sounds, or musical details beyond simple transcription. The audio component recognizes speech in 19 languages and translates between Chinese, English, and 14 others. In vision, BytePlus claims significant gains in physical and medical reasoning, fine perception, and embodied understanding.

The Agent and Coding component expands capabilities for following long instructions, task decomposition, and self-verification, with refined integration for frameworks like OpenClaw and Hermes Agent. Front-end page generation, 3D scenes, and game prototypes gain in visual quality and completeness. The GUI module, which detects buttons, menus, forms, and interface states, adds an action layer (click, input, scroll, drag & drop), making the model usable for end-to-end Browser Use and Computer Use workflows.

On public benchmarks, Dola Seed 2.0 Lite claims SOTA or near-SOTA scores: BabyVision 64.7%, VideoMME 89.0, LongVideoBench 79.0, OmniVideoBench 61.7, MedXpertQA-MM 79.6, HiPhO 83.8, GPQA Diamond 88.4%. The model is accessible via the BytePlus ModelArk platform with a context window of 262,000 tokens, at around $0.25 per million input tokens and $2 per million output tokens.