GPT-4o mini
Vision-language model Generally availableOpenAI lightweight multimodal model optimized for cost efficiency, low latency, and high-throughput text and vision workloads.
Overview
GPT-4o Mini is OpenAI's small, cost-efficient multimodal model launched in July 2024 to replace GPT-3.5 Turbo for developer API applications. Engineered to minimize latency and token costs without sacrificing core intelligence, GPT-4o Mini makes advanced text generation, vision understanding, and structured data processing accessible for high-volume enterprise pipelines.
Despite its compact footprint, GPT-4o Mini maintains a full 128,000-token context window and supports up to 16,384 output tokens per request. It accepts text and visual inputs, outperforming prior small-tier models on benchmarks such as MMLU, MGSM, and HumanEval while operating at a fraction of the inference cost.
GPT-4o Mini supports native function calling, JSON mode, structured output schema enforcement, and streaming completions. It is deployed exclusively as a hosted proprietary model via OpenAI's Direct API and cloud partner platforms. Open weights and local execution are not supported.
Technical Specifications
Context Window
128,000 tokens
Max Input Tokens
128,000
Max Output Tokens
16,384
Architecture
Dense Multimodal Transformer
Released
2024-07-18
Modalities
Text → Text (Both)
Image → Text (Input)
Capabilities
Text Generation
Code Generation
Structured Output
Image Understanding
Function Calling
Long Context
Multilingual Support
Use Cases
Content Creation
Coding Assistance
Customer Support
Languages
English (en)
Spanish (es)
French (fr)
German (de)
Chinese (zh)
Japanese (ja)
Access Routes
| Route | Type | Protocol | Region / Scope | Provider | Status |
|---|---|---|---|---|---|
| OpenAI API | Hosted API (Direct provider) | OpenAI-compatible REST API | Global | OpenAI | Available |
Identifiers
- gpt-4o-mini Canonical identifier Default
- gpt-4o-mini-2024-07-18 Canonical identifier Default