chat-completions-multimodal-input · EN · 2026-10-06

Sending Images and Multi-Part Messages Through an OpenAI-Compatible /chat/completions Call

Learn how to structure multi-part messages with images for OpenAI-compatible /chat/completions calls. Understand the content array format, model support for vision, and practical tips for working with different providers through a single API.

Introduction to Multi-Part Messages

Modern chat models can process more than just text. Many support images, and some handle other media types like audio or documents. To use these capabilities through an OpenAI-compatible /chat/completions endpoint, you need to structure your request with a content array instead of a simple string.

This guide explains how to format multi-part messages, which models accept which media types, and how to handle variations across providers.

The Content Array Format

In the OpenAI API, the content field of a message can be either a string or an array of content parts. For multi-part messages, you use an array where each element describes a part of the message.

Basic Structure

{
  "model": "your-model",
  "messages": [
    {
      "role": "user",
      "content": [
        { "type": "text", "text": "What is in this image?" },
        { "type": "image_url", "image_url": { "url": "https://example.com/image.jpg" } }
      ]
    }
  ]
}

Content Part Types

  • Text: { "type": "text", "text": "..." }
  • Image URL: { "type": "imageurl", "imageurl": { "url": "..." } }
  • Base64 Image: { "type": "imageurl", "imageurl": { "url": "data:image/jpeg;base64,..." } }

Some providers may support additional types like input_audio or file, but support varies.

Model Support for Vision

Not all models accept images. Here's a general overview of which models typically support vision inputs when accessed through the aggregator:

  • Claude models: Claude 3 and later versions support vision. They accept images via URL or base64. Supported formats include JPEG, PNG, GIF, and WebP.
  • GPT models: GPT-4 series and later support vision. They accept images via URL or base64. Supported formats: PNG, JPEG, WEBP, non-animated GIF.
  • Qwen models: Some Qwen models (like Qwen-VL) support vision. Check the specific model documentation for details.
  • Other models: DeepSeek, GLM, Kimi, etc., may have vision-capable variants. Always verify the model's capabilities before sending images.

Since the aggregator routes to different providers, the exact behavior may vary. The API will return an error if the model doesn't support the provided content type.

Practical Tips

Handle Provider Differences

While the aggregator aims to standardize the API, providers may have subtle differences in how they handle multi-part content. For example:

  • Some may require certain image formats.
  • Some may have size limits on images.
  • Some may interpret multiple images in a single message differently.

Always test with your target models and handle errors gracefully.

Image Encoding

You can provide images as URLs or as base64-encoded data URLs. Using URLs is simpler but requires the image to be publicly accessible. Base64 is useful for local files or private images, but increases request size.

Combining Text and Images

You can include multiple text and image parts in any order. The model processes them as a single message. Typically, you place text instructions before or after the image to guide the model's response.

Multiple Images

You can include several images in one message by adding multiple image_url parts. The model will consider all of them together. Be mindful of token usage and potential limits.

Example: Sending an Image from a URL

import requests

response = requests.post(
    "https://api.your-aggregator.com/v1/chat/completions",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    json={
        "model": "claude-3-sonnet",
        "messages": [
            {
                "role": "user",
                "content": [
                    {"type": "text", "text": "Describe this image."},
                    {"type": "image_url", "image_url": {"url": "https://example.com/cat.jpg"}}
                ]
            }
        ]
    }
)
print(response.json())

Example: Sending a Base64 Image

import base64

with open("image.jpg", "rb") as f:
    img_b64 = base64.b64encode(f.read()).decode("utf-8")

response = requests.post(
    "https://api.your-aggregator.com/v1/chat/completions",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    json={
        "model": "gpt-4-vision-preview",
        "messages": [
            {
                "role": "user",
                "content": [
                    {"type": "text", "text": "What is this?"},
                    {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{img_b64}"}}
                ]
            }
        ]
    }
)
print(response.json())

Troubleshooting

  • Error: Unsupported content type: The model may not support images. Check model capabilities.
  • Error: Invalid image format: Ensure the image is in a supported format (PNG, JPEG, etc.).
  • Error: Image too large: Reduce image dimensions or file size.
  • Timeouts: Large images or multiple images can increase processing time.

Conclusion

Multi-part messages with images are a powerful way to interact with vision-capable models through a unified API. By using the content array format, you can send text and images together to models from Claude, GPT, Qwen, and others. Always verify model support and handle errors gracefully to ensure a smooth experience.