Checked other resources
Example Code
import { BaseCallbackHandler } from "@langchain/core/callbacks/base";
import type { Serialized } from "@langchain/core/load/serializable";
import { HumanMessage } from "@langchain/core/messages";
import type { BaseMessage } from "@langchain/core/messages";
import { FakeListChatModel } from "@langchain/core/utils/testing";
class CaptureInputHandler extends BaseCallbackHandler {
name = "CaptureInputHandler";
messages: BaseMessage[] | undefined;
async handleChatModelStart(
_llm: Serialized,
messages: BaseMessage[][],
): Promise<void> {
this.messages = messages[0];
}
}
// Shape still shown in the JS multimodal docs:
// https://docs.langchain.com/oss/javascript/langchain/messages#multimodal
const message = new HumanMessage({
content: [
{ type: "text", text: "two PDFs" },
{
type: "file",
source_type: "base64",
mime_type: "application/pdf",
data: "JVBERi0xLjQK",
},
{
type: "file",
source_type: "base64",
mime_type: "application/pdf",
data: "JVBERi0xLjUK",
},
],
});
const handler = new CaptureInputHandler();
await new FakeListChatModel({ responses: ["ok"] }).invoke([message], {
callbacks: [handler],
});
console.log(JSON.stringify(handler.messages?.[0]?.content, null, 2));
Error Message and Stack Trace (if applicable)
No response
Description
- I'm trying to trace chat model inputs that contain several base64 file attachments in one
HumanMessage using the data content block shape shown in the TypeScript > Messages > Multimodal docs (type + source_type + data + mime_type). Observability tools (e.g., Langfuse) rely on _formatForTracing rewriting those blocks into OpenAI image_url data URIs so media can be detected and offloaded.
- I expect every matching base64/URL block in the message to be converted the same way before
handleChatModelStart fires (tracing copy only — the provider still receives the original message).
- Instead, only the first matching block is converted. Later blocks are left as authored
file/image content, so their base64 stays inline in the trace.
Actual output
[
{
"type": "text",
"text": "two PDFs"
},
{
"type": "image_url",
"image_url": {
"url": "data:application/pdf;base64,JVBERi0xLjQK"
}
},
{
"type": "file",
"source_type": "base64",
"mime_type": "application/pdf",
"data": "JVBERi0xLjUK"
}
]
The second attachment is never rewritten, so its base64 stays inline in the trace.
Expected output
[
{
"type": "text",
"text": "two PDFs"
},
{
"type": "image_url",
"image_url": {
"url": "data:application/pdf;base64,JVBERi0xLjQK"
}
},
{
"type": "image_url",
"image_url": {
"url": "data:application/pdf;base64,JVBERi0xLjUK"
}
}
]
Root cause
In _formatForTracing (https://github.com/langchain-ai/langchainjs/blob/main/libs/langchain-core/src/language_models/chat_models.ts#L171), the clone guard if (messageToTrace === message) also gates the conversion. After the first rewrite, messageToTrace !== message, so subsequent blocks are skipped.
Docs / shape note
The conversion path only understands the deprecated isBase64ContentBlock / isURLContentBlock helpers (source_type). Standard multimodal blocks are not converted at all:
| Shape |
Documented in |
Converted? |
{ type, source_type, data, mime_type } |
JS multimodal docs |
Yes — but only the first block |
{ type, base64, mime_type } |
Python multimodal docs |
No |
{ type, data, mimeType } |
JS ContentBlock.Multimodal.File |
No |
So today the only way to get attachments traced as data URIs is the deprecated shape — and even then only one attachment per message is rewritten.
Proposed fix
Clone once, then convert every matching block. We’ve verified this locally with a pnpm patch on @langchain/core@1.2.3.
System Info
@langchain/core: 1.2.3 (also present on main as of this report)
langchain: 1.5.3
Platform: macOS
Node: 24.18.0
pnpm: 11.15.1
Checked other resources
Example Code
Error Message and Stack Trace (if applicable)
No response
Description
HumanMessageusing the data content block shape shown in the TypeScript > Messages > Multimodal docs (type+source_type+data+mime_type). Observability tools (e.g., Langfuse) rely on_formatForTracingrewriting those blocks into OpenAIimage_urldata URIs so media can be detected and offloaded.handleChatModelStartfires (tracing copy only — the provider still receives the original message).file/imagecontent, so their base64 stays inline in the trace.Actual output
[ { "type": "text", "text": "two PDFs" }, { "type": "image_url", "image_url": { "url": "data:application/pdf;base64,JVBERi0xLjQK" } }, { "type": "file", "source_type": "base64", "mime_type": "application/pdf", "data": "JVBERi0xLjUK" } ]The second attachment is never rewritten, so its base64 stays inline in the trace.
Expected output
[ { "type": "text", "text": "two PDFs" }, { "type": "image_url", "image_url": { "url": "data:application/pdf;base64,JVBERi0xLjQK" } }, { "type": "image_url", "image_url": { "url": "data:application/pdf;base64,JVBERi0xLjUK" } } ]Root cause
In
_formatForTracing(https://github.com/langchain-ai/langchainjs/blob/main/libs/langchain-core/src/language_models/chat_models.ts#L171), the clone guardif (messageToTrace === message)also gates the conversion. After the first rewrite,messageToTrace !== message, so subsequent blocks are skipped.Docs / shape note
The conversion path only understands the deprecated
isBase64ContentBlock/isURLContentBlockhelpers (source_type). Standard multimodal blocks are not converted at all:{ type, source_type, data, mime_type }{ type, base64, mime_type }{ type, data, mimeType }ContentBlock.Multimodal.FileSo today the only way to get attachments traced as data URIs is the deprecated shape — and even then only one attachment per message is rewritten.
Proposed fix
Clone once, then convert every matching block. We’ve verified this locally with a pnpm patch on
@langchain/core@1.2.3.System Info