LanguageModel: promptStreaming() method

Secure context: This feature is available only in secure contexts (HTTPS), in some or all supporting browsers.

The promptStreaming() method of the LanguageModel interface sends input to the language model and returns a ReadableStream that delivers the model's response incrementally as it is generated.

This is useful for displaying responses to users incrementally for outputs that take a long time to complete, or for any scenario where perceived latency should be minimized. Consume the stream using for await...of or by attaching a reader via ReadableStream.getReader().

Syntax

js
promptStreaming(input)
promptStreaming(input, options)

Parameters

input

The content to prompt the model with. This is either:

options Optional

Options for creating a prompt. Properties include:

responseConstraint

An object following the structure defined by JSON Schema defining the precise format the model's output should be delivered in. When provided and omitResponseConstraintInput is false, any implementation-defined constraint-description message is included in the measurement.

omitResponseConstraintInput

A boolean; when true, the automatic constraint-description message is excluded from the measurement.

signal

An AbortSignal to cancel the operation.

Return value

A ReadableStream of String chunks. Each chunk represents a piece of the model's response as it is generated. The stream closes when generation completes.

Exceptions

Errors are surfaced as stream errors rather than as rejected promises. Consumers should handle errors using a stream's standard error-handling mechanisms.

AbortError DOMException

Thrown if the operation was cancelled via the signal option.

NotAllowedError DOMException

Thrown if usage of the method is blocked by a language-model Permissions-Policy.

NotSupportedError DOMException

Thrown if:

OperationError DOMException

Thrown if the prompt fails for any other reason not listed in the other exception types.

QuotaExceededError DOMException

Thrown if the prompt would cause the session's context usage to exceed the model's LanguageModel.contextWindow.

SyntaxError DOMException

Thrown if:

TypeError

Thrown if:

Description

The promptStreaming() method adds the provided input to the context window and generates a response. The entire response is receives incrementally as a ReadableStream.

To receive the response as one complete string, use LanguageModel.prompt() instead. To add content to the context window without generating a response, use LanguageModel.append().

Each call to promptStreaming() adds to the session's context. To branch from a given state without affecting the original session, call LanguageModel.clone().

Examples

undefined

Streaming a response to the page

This example writes out chunks from a promptStreaming() call's ReadableStream as they arrive.

js
const session = await LanguageModel.create();
const output = document.querySelector("#output");

const stream = session.promptStreaming("Write a short poem about the ocean.");

for await (const chunk of stream) {
  output.textContent += chunk;
}

See also Using the Prompt API > Complete streaming example.

Streaming with an abort signal

This example shows how to use an AbortController with promptStreaming().

js
const controller = new AbortController();
document
  .querySelector("#stop")
  .addEventListener("click", () => controller.abort());

const stream = session.promptStreaming("Tell me a long story.", {
  signal: controller.signal,
});

try {
  for await (const chunk of stream) {
    output.textContent += chunk;
  }
} catch (err) {
  if (err.name === "AbortError") {
    console.log("Streaming was stopped by the user.");
  }
}

Collecting streamed chunks into a single string

In this example, chunks from a ReadableStream are collected before the whole stream is written out.

js
const session = await LanguageModel.create();
const stream = session.promptStreaming("Explain quantum entanglement.");
const chunks = [];

for await (const chunk of stream) {
  chunks.push(chunk);
}

const fullResponse = chunks.join("");
console.log(fullResponse);

Specifications

See also