Skip to content

[Bug] Source insights (transformations) silently truncated at Esperanto default max_tokens=850 — no way to override for openai_compatible / non-reasoning models #1273

Description

@heinrichhermann

Environment: open-notebook 1.14.0 (v1-latest-single), esperanto 2.25.1, provider openai_compatible (vLLM, served model Qwen3.8, 256K context window).

Problem

Source insights (transformations) are silently cut off mid-sentence / mid-word. The truncation correlates directly with the transcript length — long sources get cut, short ones complete cleanly:

full_text (chars) insight (chars) result
13431 1228 / 2002 truncated (e.g. ends …Eine form)
16581 11333 truncated (mid list item)
325 3832 complete (ends with .)
715 3021 complete
451 2449 complete

Root cause

Transformations provision the language model via provision_langchain_model() (open_notebook/ai/provision.py) -> model.to_langchain(). No explicit max_tokens is passed anywhere in this path.

Esperanto's LanguageModel base default is max_tokens = 850 (esperanto/providers/llm/base.py:19). The OpenAI provider only skips this default cap for reasoning models (_is_reasoning_model() = name starts with o1/o3/o4/gpt-5, esperanto/providers/llm/openai.py:316). An openai_compatible model such as Qwen3.8 is not detected as reasoning, so every transformation request is sent with max_tokens=850.

There is no way to override this from Open Notebook:

  • ModelCreate / ModelResponse have no max_tokens field.
  • Credential.to_esperanto_config() only emits known keys (api_key, base_url, endpoint*, num_ctx, ...) — arbitrary keys like max_tokens are dropped.
  • No relevant env var / setting exists.

Verified live in the container:

m = await model_manager.get_default_model("transformation")   # OpenAICompatibleLanguageModel "Qwen3.8"
m.max_tokens                 # -> 850
m.to_langchain().max_tokens  # -> 850  (ChatOpenAI)

The backend model itself is not the limit — a direct request to the same endpoint with max_tokens=3000 returns a complete 2852-token answer (finish_reason=stop). So the cap is imposed solely by Open Notebook/Esperanto.

Suggested fix

Analogous to the Ask fix in #1241 (OPEN_NOTEBOOK_ASK_MAX_TOKENS): make the transformation output budget configurable — e.g. an OPEN_NOTEBOOK_TRANSFORMATION_MAX_TOKENS env var and/or a per-model max_tokens in the model/credential config — and/or stop silently falling back to Esperanto's 850 default for non-reasoning openai_compatible models.

Related

#1221, #1247, #1241 (Ask/Q&A truncation), #1200 (provider-declared tunable params), #947 (model & provider advanced config).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-designWanted, but the how isn't resolved — needs design/spec before it's ready

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions