Skip to content

Latest commit

 

History

History
348 lines (284 loc) · 18.9 KB

File metadata and controls

348 lines (284 loc) · 18.9 KB

Lazy-loading the generator's heavy dependencies (prototype)

Status: prototype. Wired for highlight.js today; sketches the same shape for Mermaid. See src/highlight.ts, src/highlight-hljs.ts, and src/highlight-lazy.test.ts.

The demo site (docs/index.html) exercises both paths live: its bundle is built with code splitting, so the highlight.js grammars are fetched as a separate chunk on the first fenced code block, and mermaid is fetched via an import map on the first ```mermaid fence. Status chips on the page show each load happening.

The problem

renderMarkdownUnsafe() (the "generator") is a synchronous string→HTML function. Its only heavy runtime dependency is highlight.js — the core plus a dozen language grammars. Because render-blocks.ts imported highlight.ts, which statically imported highlight.js, the whole grammar payload landed in every consumer's main bundle, even one that never renders a fenced code block.

Measured with esbuild (--bundle --minify --format=esm) on dist/:

Bundle Size Contains highlight.js?
Main entry (index.js), before ~164 KB yes (unconditionally)
Main entry (index.js), after ~88 KB no
highlighters/highlightjs chunk (lazy) ~71 KB yes (fetched on demand)

So ~71 KB — the grammars — now loads only when a host asks for highlighting.

The tracked number: a CI bundle-size gate (#113)

The figures above are minified-only, whole-payload sizes. The number a consumer actually pays over the wire is gzipped, and — because the peer dependencies (dompurify, highlight.js, katex, mermaid, shiki, entities) are the host's to bundle — it excludes them. Measured that way (--bundle --minify --format=esm --platform=browser, peers external, gzipped), the main entry is ~31.5 KB gzipped, and the emoji data subpath is the next largest at ~16.2 KB.

This is no longer a hand-run figure. scripts/check-bundle-size.mts (npm run size) measures the main entry and every key subpath against the committed budgets in scripts/bundle-size-budget.json, and the CI size job fails the build if any entry exceeds its budget. The emoji table — bundled, not external — is caught by that byte budget: dragging it into the core entry blows the . budget outright.

A peer dependency is trickier: because the host bundles it, esbuild keeps a leaked import 'shiki' as an ~30-byte external statement, well inside the ~5% headroom, while the consumer's real chunk silently grows ~70 KB. So the gate also reads esbuild's metafile and fails if the main entry statically imports any peer — a leaked grammar or math payload reddens the PR on the import itself, not on the bytes. (A lazy import() of a peer on its own subpath is the intended pattern and is left alone.) When an increase is intentional, bump the budget deliberately with npm run size:update and commit the JSON.

The shape: a pluggable backend, mirroring the sanitizer split

This reuses the exact pattern the sanitizer already uses (the sanitizerBackend config field + the @copse/streaming-markdown/sanitizers/dompurify subpath):

  • highlight.ts (core) carries no highlight.js code. It keeps only the cheap string work — language aliases, the KNOWN_LANGUAGES set, and fenceCodeClass — and a config-injected slot: the codeHighlighter field, read internally via getCodeHighlighter() and applied around each render by withConfig. With no highlighter configured, highlightFenceCode returns escaped plain text.
  • highlight-hljs.ts (backend) is the only module that imports highlight.js. It registers the grammars and exports highlightjsHighlighter (the value) and loadHighlightjs() (which returns that CodeHighlighter), to pass via MarkdownConfig.codeHighlighter. It lives behind the @copse/streaming-markdown/highlighters/highlightjs subpath, so a bundler drops it unless the host references that entry.

Why language resolution stays in the core

fenceCodeClass('ts') returns hljs lang-typescript before the backend loads and the identical string after. That stability matters for streaming: the core renders a code fence as plain escaped text with the final class immediately, and a later re-render (once the grammar chunk arrives) only swaps the interior to token spans — the <pre><code class> element never churns. KNOWN_LANGUAGES in the core must stay in sync with the grammars the backend registers.

Using it

Eager (highlighting from first paint — pulls the chunk into your bundle):

import { renderMarkdown } from '@copse/streaming-markdown'
import { highlightjsHighlighter } from '@copse/streaming-markdown/highlighters/highlightjs'

renderMarkdown(md, { codeHighlighter: highlightjsHighlighter })

Lazy (the grammars are a separate chunk, fetched only when this runs):

// e.g. on the first fenced block seen, or during an idle callback
const { loadHighlightjs } = await import('@copse/streaming-markdown/highlighters/highlightjs')
const codeHighlighter = await loadHighlightjs() // returns the highlighter

// re-render the message with { codeHighlighter } in config so already-rendered
// fences upgrade from plain → highlighted
rerender({ codeHighlighter })

Until either runs, code fences render as safe, escaped plain text with the correct hljs lang-* class.

Shiki — same registry, an async-loading backend

@copse/streaming-markdown/highlighters/shiki is a second highlighter backend for the same codeHighlighter config slot. It sits between the two patterns above: like mermaid, shiki is an optional peer dependency reached only through variable-specifier dynamic imports (the package builds without it, and zero shiki bytes can land in the main entry — or in the subpath chunk itself); like highlight.js, the backend highlights synchronously once ready.

Because shiki can only initialize asynchronously, loadShiki() is the load seam: it awaits shiki's fine-grained core (shiki/core + createHighlighterCore with the JavaScript regex engine — no oniguruma WASM, no bundled-registry entry) plus the grammar and theme modules, then registers a backend that highlights synchronously against the loaded instance. Until it resolves, fences render as escaped plain text with the same stable core-resolved hljs lang-* class, and a re-render upgrades them in place — identical UX to lazy highlight.js. Token colors are emitted as classes (not shiki's inline style attributes, which the sink sanitizer strips); see the Shiki section of EXTENDING.md for the styling contract and shikiThemeCss().

Mermaid is not bundled by this package: mermaid-source.ts is pure string preparation, and the generator only emits inert scaffolding (<div class="mermaid-diagram mermaid-diagram--pending"><pre class="mermaid">…). The heavy mermaid library is host-injected and rendered after sanitization, so it is already lazy by construction. What the package was missing is an official hook so every host stops hand-rolling the "find pending diagrams, load mermaid, inject SVG, retry on the aggressive source candidate" dance.

The same registry shape as highlighting now ships for it:

  • mermaid.ts (core) carries no mermaid code. It holds the async renderer seam (the diagramRenderer config field, consumed by hydrate() / hydratePendingDiagrams's renderer option — not the synchronous render), the DiagramRenderer interface, and hydratePendingDiagrams(root, opts?) — which finds every mermaid-diagram--pending container under root, tries the gentle then aggressive mermaidSourceCandidates() until one renders, and flips the container to --rendered (SVG injected) or --error.
  • mermaid-mermaidjs.ts (backend) is the only module that references mermaid. It lives behind @copse/streaming-markdown/diagrams/mermaid and exports mermaidDiagramRenderer and loadMermaid() (which returns the renderer value — pass it via MarkdownConfig.diagramRenderer or the hydrate renderer option). mermaid is an optional peer dependency — the host installs it, the package never bundles it. The dynamic import uses a variable specifier so the package builds and type-checks even when the peer isn't installed.
import { hydratePendingDiagrams } from '@copse/streaming-markdown'

const { loadMermaid } = await import('@copse/streaming-markdown/diagrams/mermaid')
const renderer = await loadMermaid()    // returns the backend; library loads lazily
await hydratePendingDiagrams(messageEl, { renderer }) // pending → rendered SVG

Trust boundary: mermaid SVG is produced by the trusted library after the sink sanitizer runs and is injected without re-sanitization (matching the existing design invariant). A safety-conscious host can pass hydratePendingDiagrams(root, { transformSvg }) to run the SVG through its own sanitizer first.

The mechanism is covered by src/mermaid-lazy.test.ts with a stub renderer (the real mermaid library needs a browser DOM it can't get in jsdom); the backend module is the thin adapter over mermaid.render.

KaTeX — the same shape again, for math

Math (#70) mirrors mermaid exactly. The generator emits inert scaffolding for every math form — ```math fences, $$ … $$ / \[ … \] display blocks, and $…$ / \(…\) inline spans (math-block--pending / math-inline--pending, escaped TeX source inside) — so the KaTeX payload is never needed at render time. The prose delimiter grammar (everything except the always-on ```math fence) is itself opt-in via mathSyntax (#78): until { mathSyntax: true } turns it on (the null default defers to a process-wide renderer registration), $…$-style text stays ordinary prose and output is byte-identical to a math-free build, so hosts that never opt in pay nothing at all:

  • math.ts (core) carries no KaTeX code: the async renderer seam (the mathRenderer config field, consumed by hydrate() / hydratePendingMath's renderer option — not the synchronous render), the MathRenderer interface, and hydratePendingMath(root, opts?) — which finds every pending block/span under root, renders it (display mode for blocks, inline for spans), and flips the element to --rendered (HTML injected) or --error (escaped source kept visible).
  • math-katex.ts (backend) is the only module that references katex. It lives behind @copse/streaming-markdown/math/katex and exports katexMathRenderer and loadKatex() (which returns the renderer value — pass it via MarkdownConfig.mathRenderer, with mathSyntax: true for the prose grammar, or the hydrate renderer option). katex is an optional peer dependency, imported through a variable-specifier dynamic import so the package builds and type-checks without it.
import { hydratePendingMath } from '@copse/streaming-markdown'

const { loadKatex } = await import('@copse/streaming-markdown/math/katex')
const renderer = await loadKatex()  // returns the backend; library loads lazily
// render with { mathSyntax: true } to turn the prose grammar on, then:
await hydratePendingMath(messageEl, { renderer }) // pending → rendered KaTeX HTML

The host also loads KaTeX's stylesheet/fonts (katex/dist/katex.min.css).

Trust boundary: KaTeX HTML is injected after the sink sanitizer and not re-sanitized, matching the mermaid invariant; hydratePendingMath(root, { transformHtml }) is the seam for a host that wants to (and the required hook under Trusted Types enforcement).

Emoji shortcodes — a lazy data subpath, no peer

Emoji shortcodes (#86) reuse the split for a data payload rather than a library. The optional pass and its GitHub/gemoji-aligned map live in src/emoji-shortcodes.ts / src/emoji-shortcode-map.ts behind @copse/streaming-markdown/inline/emoji; the ~50 KB alias table is pulled into a bundle only when a host imports that entry, so a consumer that never opts in pays zero bytes for it (src/emoji-shortcodes.test.ts asserts this with an esbuild bundle of the main entry). Unlike the highlighter/mermaid/KaTeX backends there is no optional peer dependency — the pass is pure string work built on the public inlinePasses config field, so it needs no registry in the core at all.

import { renderMarkdown } from '@copse/streaming-markdown'
import { emojiInlinePass } from '@copse/streaming-markdown/inline/emoji'

renderMarkdown(md, { inlinePasses: [emojiInlinePass] }) // `:smile:` → 😄; see EXTENDING.md#custom-inline-syntax-inline-passes

Input smoothing — an opt-in reveal cadence (#84)

Chunky token arrival shows as chunky updates: an LLM transport delivers text in bursts, and the incremental DOM emitter renders exactly what it is given, when it is given it. The optional input smoother steadies that reveal — it sits between the host's chunk arrival and renderer.update() and releases the growing string a few characters per frame (via requestAnimationFrame) instead of in raw token bursts.

It lives behind @copse/streaming-markdown/smoothing and is never re-exported from the main entry, so a host that doesn't import it pays zero bytes and the default emitter behaviour is byte-for-byte unchanged (the same "zero bytes in the main bundle" contract as the highlighter/mermaid/KaTeX backends above — src/smoothing.test.ts asserts an esbuild bundle of the main entry contains none of the smoother's code).

import { StreamingMarkdownRenderer } from '@copse/streaming-markdown'
import { createInputSmoother } from '@copse/streaming-markdown/smoothing'

const renderer = new StreamingMarkdownRenderer(host)
const smoother = createInputSmoother({
  update: (text) => renderer.update(text), // the sink for each revealed prefix
  cadence: 'adaptive',                     // follow the stream's own rate (default 'fixed')
})

for await (const fullTextSoFar of stream) smoother.push(fullTextSoFar)
smoother.finish(() => showFinalRender()) // stream end: drain the rest, then settle
smoother.dispose()                       // tear down (cancels any pending frame)

push(text) takes the full accumulated message so far (the same argument renderer.update takes), so every value the smoother releases is a prefix of the target. That is exactly the streaming contract the pending-state machinery and DOM morph already converge on, so smoothing composes with them for free rather than fighting the morph.

Cadence: fixed or adaptive

  • 'fixed' (the default, for compatibility) walks the revealed prefix at a constant charsPerSecond (default 600). It is predictable, but it does not follow the stream: a transport slower than the rate still shows as bursts (each chunk drains in a frame or two, then the reveal waits for the next), and one faster than the rate falls further and further behind.
  • 'adaptive' — recommended for LLM output — reveals at a velocity that tracks the arrival rate, running about lagMs (default 120) behind it. The velocity is low-pass filtered, so a steady stream reveals steadily, a burst speeds the reveal up over a few frames rather than in one jump, and the lag stays bounded however fast the model is. After a frame gap long enough to mean the page was hidden, it catches up at once instead of replaying text.

Measured on a transport delivering 12 characters every 50ms, 'fixed' leaves the text frozen in roughly two frames of three; 'adaptive' advances it in nearly every frame, 4–5 characters at a time.

Ending the stream: finish or flush

  • finish(onSettled?) reveals whatever is still pending over a short drain (~60ms of lag, whatever the cadence), then calls onSettled once — the place to swap in a final, at-rest render. It settles synchronously when nothing is pending, smoothing is off, or the page is hidden. A push after finish means the stream resumed and drops the pending callback.
  • flush() releases everything immediately (and settles a pending finish). Use it when completion must carry no delay at all.

Where a frame may end

A revealed prefix never ends right after markdown punctuation or whitespace (` * _ ~ [ ] ( ) | # > ! - + = . : < \ and digits): the cut moves forward past that run — a bounded few characters — so a frame always ends just after a word character. Otherwise the half-arrived syntax would show raw for a frame before the renderer could tell what it becomes: the first two backticks of a closing fence, the || of a table separator, a 1. before its list item, the brackets of [text] before (url) arrives, or a trailing newline briefly opening an empty line. A cut also never splits a surrogate pair.

Re-rendering mid-stream

Pass initial with the text already on screen when a host rebuilds a message element partway through a stream: the reveal starts from its end instead of replaying the whole message.

Why input smoothing, not output animation

Two designs were on the table (see #84): throttle the input string fed to update(), or animate the output (CSS/opacity transitions on newly-added nodes). We ship input smoothing only — it is lower-risk and framework- and theme-agnostic, and it reuses the renderer's existing convergence guarantees instead of introducing a second animation system that could fight the DOM morph. CSS entrance animations are deliberately left to the host: a host that wants nodes to fade in can style the emitter's class hooks (.stream-pending, .stream-complete) with its own transitions — that is a presentation concern the package does not own.

Guarantees

  • Convergence. After flush(), or once finish() settles, the sink has seen the full text, so the rendered DOM equals a single un-smoothed update(fullText) — asserted against a direct full-string render in the tests, for both cadences.
  • prefers-reduced-motion. Honoured by default: when the environment reports reduced motion, smoothing is disabled and push passes straight through immediately (respectReducedMotion: false overrides; disabled: true forces pass-through unconditionally). smoother.enabled reports which mode is active.
  • No completion lag on demand. flush() releases the whole pending target at once and stops the loop; finish() drains it within a few frames.
  • Never recursive. A frame scheduler that calls back synchronously (some test shims do) cannot pace anything; the smoother detects it and releases text as it is pushed instead of recursing.
  • Environment-guarded. requestAnimationFrame / cancelAnimationFrame, performance.now, and matchMedia are all read through defaulted, injectable seams (requestFrame / cancelFrame / now / matchMedia options), so the smoother runs — and is unit-tested with a fake clock — under node/jsdom where those globals may be absent.