DeepSeek-V2.5
General language model Generally available Open WeightsDeepSeek's 236B Mixture-of-Experts (21B active) open-weights model combining general chat and advanced code generation with a 128k context window.
Overview
DeepSeek-V2.5 is an open-weights Mixture-of-Experts (MoE) language model released by DeepSeek in September 2024. It unifies DeepSeek-V2 Chat and DeepSeek-Coder-V2 into a single upgraded model, offering strong general instruction following, reasoning, and programming capabilities.
DeepSeek-V2.5 comprises 236 billion total parameters, of which 21 billion are activated per token. It features Multi-head Latent Attention (MLA) and DeepSeekMoE architecture, supporting a 128,000-token context window with a 4,096-token maximum API output limit.
The model is open-sourced under the MIT License for commercial and research use. It is available on the official DeepSeek API (`deepseek-chat` endpoint at $0.14 / 1M input tokens, $0.28 / 1M output tokens) and can be self-hosted.
Technical Specifications
Context Window
128,000 tokens
Max Input Tokens
128,000
Max Output Tokens
4,096
Parameters
236.0B
Architecture
MoE Transformer
Released
2024-09-05
Modalities
Text → Text (Both)
Capabilities
Text Generation
Code Generation
Reasoning
Function Calling
Long Context
Multilingual Support
Use Cases
Coding Assistance
Data Analysis
Research
Languages
English (en)
Chinese (zh)
Japanese (ja)
Access Routes
| Route | Type | Protocol | Region / Scope | Provider | Status |
|---|---|---|---|---|---|
| DeepSeek | Hosted API (Direct provider) | OpenAI-compatible REST API | Global | DeepSeek | Available |
Identifiers
- deepseek-chat Canonical identifier Default
- deepseek-chat Canonical identifier Default
- deepseek-v2.5 Canonical identifier Default