Give a text-only model the ability to work with images, with no core changes. A filter takes the image out of the request (so the text-only model never 404s on an image it cannot accept) and leaves a marker in its place. A tool then lets the model send that image to a separate vision model on demand, asking whatever it wants, as many times as it wants. The image itself stays in the chat untouched.
Important
Requires Open WebUI 0.11.0 or newer. Both parts resolve chat ids through the core helper added in that release. They will not load on older versions.
Tip
🚀 Jump to Setup — install both parts, set two valves, done in about a minute.
Note
The filter and the tool are a pair. The filter keeps the image out of the text-only model's request; the tool lets that model look at the image on demand through a vision model you configure. Install both for the on-demand flow.
Route an image to a text-only model and the request fails: the model (or the provider) rejects an image_url part it was never built to accept. The usual workaround is to hard-swap the image for a one-shot description, which throws the real image away and locks you into whatever that single description happened to capture.
Vision Bridge keeps the image and defers the looking. The text-only model drives its own conversation and, whenever it needs to, calls out to a vision model with a specific question. Ask again later with a different question and it looks again, at the same untouched image.
- Filter (
strip_only) runs on the request to the text-only model. Each image part is replaced with a text marker:[Image attached — file_id: <id>. Call analyze_image(...) to inspect it.], or[Image attached. Call analyze_image(query="…") to inspect the most recent image.]when the id is not in the request (Open WebUI inlines uploaded images before filters run). The model receives the marker, never the image. The image stays in the chat and in storage. - Model calls
analyze_image(file_id, query)whenever it needs to see something. The tool resolves the file id to the stored image, sends it plus the question to the configured vision model, and returns the answer as text. - Re-query any time. Because the image is never consumed or deleted, the model can call the tool again with a new question and get a fresh, different answer about the same image.
┌──────────────┐ image stripped ┌──────────────┐
│ Text-only │◀───to a marker──────│ Vision │
│ model │ │ Bridge │
│ (deepseek…) │──analyze_image()───▶│ Filter+Tool │
└──────────────┘◀──answer as text────└──────┬───────┘
│ file_id -> image
▼
┌──────────────┐
│ Vision model │
│ (gpt-4o, │
│ minimax…) │
└──────────────┘
| File | Type | Install location |
|---|---|---|
filter.py |
Filter | Admin Panel → Functions |
tool.py |
Tool | Workspace → Tools |
- Copy the contents of
filter.py. - In Open WebUI, go to Admin Panel → Functions → + New, paste the code, and Save.
- Enable it on your text-only model (or Global), and set the valve
strip_only = true.
- Copy the contents of
tool.py. - Go to Workspace → Tools → + Create New, paste the code, and Save.
- Open the tool's valves and set
vision_model_idto your multimodal model (e.g.gpt-4o, or a vision model on OpenRouter).
- Admin Panel → Settings → Models, edit your text-only model.
- Under Tools, enable Vision Bridge, and Save.
That is it. Send an image in a chat with that model: the filter strips it to a marker, and the model calls analyze_image when it wants to look.
| Valve | Default | Purpose |
|---|---|---|
strip_only |
true |
Remove images from the request, replacing each with a text marker, and leave them in the chat. Pair with the tool for on-demand re-analysis. Set false for describe mode (below). |
vision_model_id |
"" |
Vision model used only in describe mode. |
analysis_prompt |
"Describe this image…" | Instruction sent to the vision model in describe mode. |
label |
"Image description" | Heading for the inlined description (describe mode). |
purge_from_history |
true |
Describe mode: replace the saved image with its description. |
delete_file_record |
true |
Describe mode: delete the image file after analysis. |
skip_if_vision_capable |
false |
If the target model already has the vision capability, do nothing (useful for a Global install). |
max_images |
4 |
Max images per message (describe mode). |
| Valve | Default | Purpose |
|---|---|---|
vision_model_id |
"" |
Required. The vision-capable model that actually looks at images. |
default_query |
"Describe this image…" | Question used when the model calls analyze_image without one. |
The filter has two modes, chosen by the strip_only valve:
strip_only = true(default, tool-driven): the recommended pairing. The image is swapped for a marker and kept in the chat, and the tool does the looking on demand. Best for models that can tool-call, and the only mode that supports re-querying the same image with new questions over time.strip_only = false(describe-and-replace): the filter itself runs one vision pass up front and swaps the image for the resulting text (needsvision_model_idon the filter). For models that cannot tool-call. This consumes the image (optionally deleting it), so there is no later re-analysis. Any image the vision pass does not cover (an older turn whose image is still in the chat, one that can no longer be read, or anything pastmax_images) is replaced by a "not sent to this model" note, so the text-only model never receives an image.
Verified end-to-end against OpenRouter. A text-only deepseek-v4-flash received the marker (no image), then re-queried the same image twice via minimax-m3: "what colors?" and "any text?" returned correct, different answers. Re-analysis of one image with new questions over time works, and a vision call only happens when the model actually asks.
Do I need both the filter and the tool?
For the on-demand flow, yes. The filter keeps the image out of the text-only model's request (so it does not error), and the tool is what lets the model actually look at the image when it decides to. If your model cannot tool-call, use the filter alone in describe mode (strip_only = false), which inlines one description up front.
Which model does the actual looking?
Whatever you set as vision_model_id on the tool (for the on-demand flow) or on the filter (for describe mode). It can be any vision-capable model your instance can reach, for example gpt-4o or a multimodal model on OpenRouter. The text-only model never sees the image itself; it only ever gets text back.
Can the model ask more than one question about the same image?
Yes, that is the point of strip_only mode. The image is left untouched in the chat, so the model can call analyze_image again with a new query at any time and get a fresh answer. Note that uploaded images reach the filter without a file id (Open WebUI inlines them first), so the tool resolves the most recent image in the chat; an older image in a multi-image chat can only be targeted when the marker carries a file_id. Describe mode does not support re-analysis at all, since it consumes the image up front.
What happens to the original image?
In strip_only mode it stays in the chat and in storage, unchanged. In describe mode it is replaced in history by its text description, and (with the default valves) the file record is deleted after analysis.
- 1.0.1 — Fixed images reaching the text-only model. Describe mode only replaced the newest image message, so images from earlier turns were sent as-is (Ollama
500 image input is not supported, or a confident description of an image the model cannot see); anything the vision pass does not cover is now replaced by a marker in both modes. Describe mode also purges the analyzed image from the chat again: Open WebUI inlines uploaded images as data URIs before filters run, so the file id is now taken from the stored chat instead of the request url, which also restores a usableanalyze_imagehint instrip_onlymode. The tool now recognises thetemporary:chat id prefix added in 0.11.0, so it no longer looks a temporary chat up in the database. - 1.0.0 — Initial release.
strip_onlytool-driven mode: the image is kept in the chat and inspected on demand viaanalyze_image, so it can be re-queried with new questions. Describe-and-replace mode is available for models that cannot tool-call.