vision-image-input · EN · 2026-09-27

Image / vision input

Learn how to send image inputs to multimodal models via the OpenAI-compatible format using image_url content blocks, including base64 encoding and URL references.

How image inputs work in the OpenAI-compatible format

Many multimodal models accept images alongside text in a single chat completion request. The OpenAI-compatible format uses a content array where each item represents a part of the message: either text or an image. For images, you use a content block with type: "image_url" and provide the image data as a URL or a base64-encoded data URI.

Building an image_url content block

To include an image, structure the message with a content array containing one or more objects. Each image object has a type field set to "imageurl" and an imageurl field. That field can be either a direct URL to the image or a base64 data URI.

{
  "model": "your-model-name",
  "messages": [
    {
      "role": "user",
      "content": [
        { "type": "text", "text": "What is in this image?" },
        {
          "type": "image_url",
          "image_url": { "url": "https://example.com/image.jpg" }
        }
      ]
    }
  ]
}

Using base64-encoded images

If the image is not publicly accessible or you prefer not to host it, you can embed it as a base64 data URI. The image_url.url value should be a string in the format data:image/<format>;base64,<data>. For example, for a JPEG image:

{
  "type": "image_url",
  "image_url": {
    "url": "data:image/jpeg;base64,/9j/4AAQSkZJRg..."
  }
}

Base64 encoding increases payload size, so use it for smaller images or when necessary. Most HTTP clients can handle base64 encoding automatically if you read the image file and encode it.

Combining multiple images and text

You can include multiple image blocks and text blocks in any order. This allows the model to compare images, answer questions about a sequence, or follow instructions that reference several images.

{
  "role": "user",
  "content": [
    { "type": "text", "text": "Compare these two charts." },
    { "type": "image_url", "image_url": { "url": "https://example.com/chart1.png" } },
    { "type": "image_url", "image_url": { "url": "https://example.com/chart2.png" } }
  ]
}

Model support and considerations

Not all models accept image inputs. Check your provider's documentation to confirm multimodality. When using an LLM API aggregator that routes to multiple models, ensure the selected model supports vision tasks. The request format remains consistent, but the underlying model must have vision capabilities.

Image size and resolution can affect processing time and cost. Some providers may have limits on image dimensions or file size. Refer to the specific model's guidelines for details.

Practical tips

  • Prefer URLs for publicly accessible images to reduce request size.
  • Use base64 only when necessary, and consider compressing images to reduce payload.
  • Place the most important image first if the model prioritizes early content.
  • Test with a simple prompt to confirm the model interprets the image as expected.
  • Remember that billing is based on the model's official pricing, with your aggregator's multiplier applied.