DeepSeek-V4-Pro and V4-Flash join BytePlus ModelArk

BytePlus ModelArk integrates DeepSeek-V4-Pro and V4-Flash, offering a one-million-token context and optimized cache efficiency for agentic coding workflows.

The two DeepSeek V4 variants are joining ModelArk, BytePlus's model platform, which takes the opportunity to advocate for a cost perspective beyond the simple price per token.

In agentic and long-context use cases, where input tokens significantly exceed output tokens and where the same corpus (codebase, internal documentation, knowledge repositories) is repeatedly accessed, cache efficiency, according to them, weighs more heavily on the price-performance ratio than the unit price, alongside throughput, service stability, and usage limits.

Pro and Flash split the roles: the former targets demanding reasoning and agentic workflows, while the latter prioritizes the balance between code performance, speed, and efficiency for daily development. On ModelArk, these open-source models come with a one-million-token context, limits up to 15,000 requests and 1.5 million tokens per minute, API access, a Playground for initial testing, and BytePlus's Coding Plan, which groups several models for assisted coding.