Skip to content

In-memory rolling window for images to reduce 413 and too-many-images errors #48095

Description

@Sugamsss

Problem

When agent loops use browser tools (Playwright, Chrome DevTools) or iterate on
visual UI tasks, a chat can accumulate many screenshot attachments. Eventually
the provider or gateway may reject the request with errors such as:

  • 413 Request Entity Too Large
  • Request contains too many images
  • provider-specific body-size limit errors

Once this happens, sending a plain text message can remain difficult because
the accumulated image payload is assembled again. Compaction can also be
affected if it receives the same raw image data.

Proposed Solution

Add an in-memory rolling window for multimodal image parts during context
assembly, before dispatch to the provider:

  1. Active image window: Keep a configurable number of recent images (for
    example, the latest 7) as raw image payloads so recent work keeps full visual
    fidelity.
  2. Payload budget: Apply a configurable cumulative image-payload budget in
    addition to the count limit. The community implementation uses a default of
    16 MiB estimated wire Base64 bytes.
  3. Older images become context cards: Replace images outside either budget
    with lightweight text cards containing the file path, what was observed,
    why it was captured, and a way to request the image again.
  4. Natural recall: If a user asks about an older screenshot, the model can
    read the path from the card and bring that image back into active context.
  5. History untouched: Only the outgoing context is changed. Persisted
    conversation history remains complete.
  6. Compaction safety: During compaction, pass text cards instead of raw
    image data.

Benefits

  • Reduces the chance of 413 and image-count failures during long visual
    debugging workflows.
  • Limits image request bandwidth and input-token pressure.
  • Keeps recent images available at full fidelity while retaining a useful,
    best-effort summary of older images.
  • Does not require a model call or an extra user prompt to prune context.

The community implementation is available at
Sugamsss/opencode-prune-images.
It is source-only and currently not published to npm. Its tests use synthetic
fixtures; provider-specific request limits and native core integration still
need independent verification.

If the core team wants this behavior natively, I would be happy to help adapt
the design to OpenCode's context assembly and compaction paths.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions