Skip to main content
Embed text, images, or one interleaved text+image record into a shared space. Exactly one of three modes is selected by which arguments are present:
  • text only — one row per string.
  • image_input only — one row per image.
  • both — a single interleaved row; text must be one string whose <|image|> placeholders are filled from image_input in order.

Parameters

List[str] | None
Up to 64 UTF-8 strings. Required unless image_input is given.
Union[ImageInput, Sequence[ImageInput], None]
One image or up to 64 images (path, URL, PIL Image, or numpy array). Required unless text is given; capped at 8 in interleaved mode.
str | None
Task instruction prefix applied to text only — e.g. "SearchQuery" for queries and "Document" for corpus items. Omitting it still works but reduces retrieval precision.
int | None
Matryoshka output width: 768 (default), 512, 256 or 128. Rows are re-normalized after truncation, and queries must share a dimension with the corpus they are scored against.
float | None
Optional timeout in seconds for the HTTP request.

Returns

Dict[str, Any]: Dict with features of shape (N, truncate_dim), float32, L2-normalized per row (cosine similarity == dot product); mode ("text", "image" or "interleaved"); model_revision of the backing checkpoint; and normalized=True.

Example