minimax_api_key enables all capabilities.
Text Chat
Image Understanding
MiniMax-M3 natively accepts text, image, and video input through the OpenAI-compatible chat endpoint. Onceminimax_api_key is configured, the Agent’s Vision tool automatically routes image requests to the selected MiniMax-M3 model, with no need to specify a separate vision model.
Image Generation
image-01.
Text-to-Speech (TTS)
Common voice examples:
For the full voice list (70+ voices across Chinese / Cantonese / English / Japanese / Korean), see the system voice list, or select visually in the Web Console under “Model Management → Text-to-Speech”.
