bytedance

ByteDance Models

Browse models from ByteDance
Models · 16
6.07Mtokens
-Cache Hit Rate

Doubao ASR is ByteDance/Volcengine's large-model speech recognition service for transcribing recorded audio files into text. Built on a 2-billion-parameter audio encoder trained on massive real-world speech data, it delivers high-accuracy transcription across noisy, multi-speaker, and domain-specific recordings.

Key capabilities:

  • High accuracy — ~30% lower error rate than traditional models, and 50%+ lower in vertical domains (music, tech, education, medical). Accent error rate down ~60%, noise/background-speech errors down 30–50%.
  • Multilingual & dialects — Chinese plus dialects (Shanghainese, Min Nan, etc.); the 2.0 model (Doubao-Seed-ASR-2.0) adds 13 foreign languages (Japanese, Korean, German, French, …) and context-aware inference for proper nouns, names, and homophones.
  • Rich text formatting — built-in automatic punctuation, inverse text normalization (numbers), semantic smoothing, and intelligent sentence segmentation — all individually toggleable.
  • Long-form audio — handles files up to 5 hours.
Input Type
Output Type
Input$0.0000327/seconds
Output-
Context-
Max Output-
16.05Mtokens
86.58%Cache Hit Rate

Doubao-Seed-Code has been deeply optimized for Agentic Programming tasks, delivering exceptional performance across multiple authoritative benchmarks — including Terminal Bench, SWE-Bench-Verified-Openhands, and Multi-SWE-Bench-Flash-Openhands — outperforming domestic counterparts and supporting a context window of up to 256k tokens.

Input Type
Output Type
Input$0.17-0.39/M tokens
Output$1.12-2.25/M tokens
Context-
Max Output-
203.32Mtokens
0%Cache Hit Rate

An all-new model purpose-built and optimized for multimodal agent scenarios. Stronger agent capabilities, upgraded multimodal understanding, and more flexible context management

Input Type
Output Type
Input$0.11-0.34/M tokens
Output$0.28-3.41/M tokens
Context-
Max Output-
1.01Btokens
24.67%Cache Hit Rate

Doubao-Seed-2.0-mini is designed for low-latency, high-concurrency, and cost-sensitive scenarios, emphasizing rapid response and flexible inference deployment. Its model performance is comparable to Doubao-Seed-1.6. It supports 256k context, four levels of think length, and multimodal understanding, making it suitable for lightweight tasks where cost and speed are paramount.

Input Type
Output Type
Input$0.03-0.12/M tokens
Output$0.28-1.12/M tokens
Context-
Max Output-
2.09Btokens
14.13%Cache Hit Rate

Doubao-Seed-2.0-lite is a balanced model designed for high-frequency enterprise scenarios, balancing performance and cost, and its overall capabilities surpass those of its predecessor, Doubao-Seed-1.8. It excels in production-oriented tasks such as unstructured information processing, content creation, search and recommendation, and data analysis, supporting long contexts, multi-source information fusion, multi-step instruction execution, and high-fidelity structured output. It significantly optimizes costs while ensuring stable performance.

Input Type
Output Type
Input$0.09-0.25/M tokens
Output$0.51-1.51/M tokens
Context-
Max Output-
238.78Mtokens
65.35%Cache Hit Rate

Doubao-Seed-2.0-Code is optimized for enterprise-level programming needs. Building on the excellent Agent and VLM capabilities of Seed 2.0, it has significantly enhanced code capabilities. Not only does it have outstanding front-end performance, but it has also been specially optimized for the multi-language coding requirements common in enterprises, making it suitable for integration with various AI programming tools.

Input Type
Output Type
Input$0.45-1.34/M tokens
Output$2.24-6.71/M tokens
Context-
Max Output-
996.90Mtokens
19.35%Cache Hit Rate

Doubao-Seed-2.0-pro is a flagship, all-around general-purpose model designed for complex reasoning and long-chain task execution scenarios in the Agent era. It emphasizes multimodal understanding, long-context reasoning, structured generation, and tool-enhanced execution. Its capabilities in executing complex instructions and multiple constraints are outstanding, stably handling scenarios such as multi-step complex planning, complex graph and text reasoning, video content understanding, and high-difficulty analysis.

Input Type
Output Type
Input$0.45-1.34/M tokens
Output$2.24-6.71/M tokens
Context-
Max Output-
5.45Btokens
76.53%Cache Hit Rate

Doubao-Seed-2.1 is a next-generation large model built for the era of coding and agents. It offers two versions, Pro and Turbo, respectively designed for exploring highly complex tasks and for large-scale production scenarios. Compared with the previous generation, Seed 2.1 delivers comprehensive upgrades across three key areas: coding engineering delivery, long-horizon agent task execution, and multimodal understanding. With stronger autonomous planning and dynamic self-repair capabilities, it is well suited to high-value production tasks in real enterprise R&D as well as broader economic and social contexts. For developers, it also offers an Evolving version for coding, office work, and productivity enhancement scenarios, with weekly updates and continuous evolution.

Input Type
Output Type
Input$0.4424/M tokens
Output$2.212/M tokens
Context-
Max Output-
38.20Btokens
70.13%Cache Hit Rate
20% OFF

Doubao-Seed-2.1 is a next-generation large model built for the era of coding and agents. It offers two versions, Pro and Turbo, respectively designed for exploring highly complex tasks and for large-scale production scenarios. Compared with the previous generation, Seed 2.1 delivers comprehensive upgrades across three key areas: coding engineering delivery, long-horizon agent task execution, and multimodal understanding. With stronger autonomous planning and dynamic self-repair capabilities, it is well suited to high-value production tasks in real enterprise R&D as well as broader economic and social contexts. For developers, it also offers an Evolving version for coding, office work, and productivity enhancement scenarios, with weekly updates and continuous evolution.

Input Type
Output Type
Input$0.70784/M tokens
Output$3.5392/M tokens
Context-
Max Output-
431.73Mtokens
87.33%Cache Hit Rate

Doubao-Seed-2.1 is a next-generation large model built for the era of coding and agents. It offers two versions, Pro and Turbo, respectively designed for exploring highly complex tasks and for large-scale production scenarios. Compared with the previous generation, Seed 2.1 delivers comprehensive upgrades across three key areas: coding engineering delivery, long-horizon agent task execution, and multimodal understanding. With stronger autonomous planning and dynamic self-repair capabilities, it is well suited to high-value production tasks in real enterprise R&D as well as broader economic and social contexts. For developers, it also offers an Evolving version for coding, office work, and productivity enhancement scenarios, with weekly updates and continuous evolution.

Input Type
Output Type
Input$0.8848/M tokens
Output$4.424/M tokens
Context-
Max Output-
81.38Btokens
37.85%Cache Hit Rate

Doubao-Seed-Character is a high-efficiency open-source large language model developed by ByteDance’s Doubao team, purpose-built for immersive roleplay interaction scenarios. It supports fast, high-fidelity alignment with diverse personas including original custom roles, fictional and real-world figures, delivering consistent, vivid, natural character-driven conversations with low deployment thresholds.

Input Type
Output Type
Input$0.1179-0.1768/M tokens
Output$0.2947-0.8842/M tokens
Context-
Max Output-
ByteDance: Doubao-Seedream-5.0-pro

ByteDance: Doubao-Seedream-5.0-pro

bytedance/doubao-seedream-5.0-pro
16.51Mtokens
-Cache Hit Rate

Seedream-5.0-pro is ByteDance’s latest image generation model, advancing image creation into a new stage of controllable production. This update brings more controllable editing, more production-ready workflows, and more natural visual results.

Input Type
Output Type
Input$0.003/counts
Output$0.0443-0.0885/counts
Context-
Max Output-
ByteDance: Doubao-Seedream-5.0-lite

ByteDance: Doubao-Seedream-5.0-lite

bytedance/doubao-seedream-5.0-lite
27.83Mtokens
-Cache Hit Rate

“Doubao-Seedream-5.0-lite is ByteDance’s latest image-generation model. For the first time, it integrates an online retrieval feature, enabling it to incorporate real-time web information and improve the timeliness of generated images. At the same time, the model’s intelligence has been further upgraded, allowing it to accurately interpret complex instructions and visual content. In addition, the model has been enhanced in terms of the breadth of world knowledge, reference consistency, and generation quality in professional scenarios, better meeting enterprise-level visual creation needs.”

Input Type
Output Type
Input-
Output$0.032/counts
Context-
Max Output-
3.72Btokens
-Cache Hit Rate

Seedance 2.5 is a new-generation multimodal video creation model developed by the Doubao foundation model team. It marks an industrial-grade step forward in long-form storytelling, reference consistency, precise editing, and multilingual video generation.

Input Type
Output Type
Input-
Output$6.4-11.7/M tokens
Context-
Max Output-
13.14Btokens
-Cache Hit Rate

Seedance 2.0 delivers ultra-realistic, highly stable audiovisual output: with outstanding motion stability and fine visual detail, it produces imagery with a live-action-quality look and an almost indistinguishable blend of real and virtual visual impact. It can handle complex scenes with ease—vividly recreating everything from subtle micro-expressions and intense physical confrontations to dynamic, high-energy song-and-dance performances. It also comes with professional camera movement, multi-shot storytelling, and text-to-video generation capabilities, enhancing narrative tension. Audio and visuals are generated natively in sync, accurately matching the visuals with rich sound effects. It supports performances in multiple languages, accents, and dialects, giving videos exceptional completeness and immersion. Seedance 2.0 is deeply optimized for three key scenarios: commercial advertising, film/TV production, and social media marketing. With industrial-grade generation quality, the hit rate for successful generations is significantly improved, lowering the barrier and cost of producing high-quality content, streamlining the workflow from idea to final cut, and delivering substantial efficiency gains for the industry.

Input Type
Output Type
Input-
Output$2.353-7.5002/M tokens
Context-
Max Output-
83.66Mtokens
-Cache Hit Rate
Sunset: 2026/09/24

Seedance 1.5 Pro is a new-generation professional-grade audio-visual co-generation video model released by the Doubao large-model team. Building on its predecessor’s multi-shot storytelling and high-definition generation capabilities, it natively supports integrated audio-and-video output, aiming to deliver an end-to-end synchronized creation experience across visuals, voice, music, and sound effects. The model also includes a built-in first-and-last-frame feature: creators only need to set the opening and ending frames of a video to precisely lock in its style, composition, and characters, which then drives the model to generate smooth, natural motion between frames. By combining audio-visual co-generation with first/last-frame control, Seedance 1.5 Pro significantly improves the efficiency, controllability, and artistic expressiveness of professional video creation.

Input Type
Output Type
Input-
Output$1.17-2.33/M tokens
Context-
Max Output-