Executable programs and extensions

Run a Local LLM in Your Browser with Transformers.js

Build a small web app that answers prompts inside your browser.

Transformers.js runs AI models from JavaScript. Once an ONNX model is downloaded, the browser can generate replies locally. This guide builds a prompt field and a reply button using LFM2.5-350M, the official example model for version 4.3.0.

Requirements and key details
  • Serve the static HTML from `localhost` or HTTPS and import Transformers.js 4.3.0 as an ES module.
  • CPU/WASM is the default; select `device: 'webgpu'` only when the browser and GPU support it.
  • The prompt is passed to a model running in the browser and does not call a separate inference server.

What changes with browser-local inference?

A typical AI website sends a prompt to a server and receives a reply. Browser-local inference1 moves that work to the user's computer. Create a Transformers.js `pipeline2()` once, then pass it a prompt: the downloaded model generates a reply inside the page.

No separate Python inference server is needed, making this useful for testing small web features. Start with lightweight tasks such as text classification or short replies. Consider a dedicated serving engine for many concurrent users or continuous large-model serving.

A laptop on a desk showing a browser window running Transformers.js with an ONNX model
Once loaded, the model processes prompts in the same browser.

Choose a model and device supported by 4.3.0

Transformers.js 4.3.0 was released on September 16, 2026. Its release example uses `q4f16` for WebGPU3 with `onnx-community/LFM2.5-350M-ONNX`. The model repository also has a `q4` file for CPU/WASM, but no `q8` file was listed, so the code below selects between the two available formats based on the device.

When switching to another ONNX model, confirm that the selected dtype file exists in that repository. The file and inference settings must match for the model to load.

Official sources do not specify minimum system RAM4 or GPU5 memory for this model. Do not infer whether it will run from the model file size alone; check the tab's memory use in the browser task manager during the first run.

These are the execution formats and browser device paths confirmed in this model repository.
Execution pathSettingWhen to use
CPU/WASMOmit `device`, use `dtype: 'q4'`Default inference path using the q4 file present in this repository
Supported GPU`device: 'webgpu'`, `dtype: 'q4f16'`When a WebGPU adapter is available and supports this format
q8Do not select for this modelNo q8 ONNX file is listed in the model repository

These are the execution formats and browser device paths confirmed in this model repository.

CPU/WASM

Setting
Omit `device`, use `dtype: 'q4'`
When to use
Default inference path using the q4 file present in this repository

Supported GPU

Setting
`device: 'webgpu'`, `dtype: 'q4f16'`
When to use
When a WebGPU adapter is available and supports this format

q8

Setting
Do not select for this model
When to use
No q8 ONNX file is listed in the model repository
A laptop beside a workstation, each representing CPU and GPU execution paths
Use CPU/WASM when WebGPU is unavailable. Processing speed and memory use vary by device.

Prepare a static HTML page on localhost

You only need two static files to start. Because the page uses an ES module CDN import, serve it with a local web server instead of opening the file directly from a file browser. `localhost` is treated as a secure context for WebGPU. Use HTTPS for a public deployment.

In a new empty folder, create `index.html` and `main.js`, then save the HTML below. This follows the official documented approach of importing the library as a module from jsDelivr. Pinning the version in the URL makes this example request 4.3.0 even if a later CDN tag changes.

index.html — input and output elements
<!doctype html>
<html lang="en">
  <meta charset="utf-8">
  <meta name="viewport" content="width=device-width, initial-scale=1">
  <title>Browser-local text generation</title>
  <label>Prompt <input id="prompt" value="Reply in one short sentence: local AI runs in a browser."></label>
  <button id="run" disabled>Generate</button>
  <pre id="status">Loading model…</pre>
  <pre id="output"></pre>
  <script type="module" src="./main.js"></script>
</html>
The ES module script loads `main.js` from the same folder.
A laptop screen with the browser downloads list and a folder containing model files
Initial network requests fetch the library and weight files; prompt generation then runs in the page.

Load the model and generate once

The following `main.js` checks for WebGPU before loading the model. It selects `q4f16` when a GPU adapter is available, or CPU/WASM `q4` if no adapter is available or the check fails. The Generate button becomes enabled when the model is ready.

On the first run, the library is fetched from a CDN and model files from Hugging Face Hub. This example processes prompts in the browser rather than sending them to a separate inference server. When integrating it into a service, also check whether analytics or logging code transmits input. Load the model once and reuse the same pipeline for each click.

main.js — check WebGPU, then run a local pipeline
import { pipeline } from 'https://cdn.jsdelivr.net/npm/@huggingface/transformers@4.3.0';

const status = document.querySelector('#status');
const output = document.querySelector('#output');
const promptInput = document.querySelector('#prompt');
const button = document.querySelector('#run');
const modelId = 'onnx-community/LFM2.5-350M-ONNX';
let adapter = null;
try {
  adapter = await navigator.gpu?.requestAdapter() ?? null;
} catch (error) {
  console.warn('WebGPU unavailable; using CPU/WASM.', error);
}
const device = adapter ? 'webgpu' : undefined;
const dtype = adapter ? 'q4f16' : 'q4';
let generator;

try {
  status.textContent = `Loading ${modelId} (${device ?? 'CPU/WASM'}, ${dtype})…`;
  generator = await pipeline('text-generation', modelId, { dtype, ...(device ? { device } : {}) });
  status.textContent = `Ready: ${device ?? 'CPU/WASM'} / ${dtype}`;
  button.disabled = false;
} catch (error) {
  status.textContent = `Model load failed: ${error.message}`;
  console.error(error);
}

button.addEventListener('click', async () => {
  if (!generator) return;
  button.disabled = true;
  status.textContent = 'Generating…';
  try {
    const result = await generator(promptInput.value, { max_new_tokens: 48 });
    const generated = result[0]?.generated_text;
    output.textContent = typeof generated === 'string'
      ? generated
      : Array.isArray(generated)
        ? generated.at(-1)?.content ?? JSON.stringify(result)
        : JSON.stringify(result);
    status.textContent = 'Done.';
  } catch (error) {
    status.textContent = `Generation failed: ${error.message}`;
    console.error(error);
  } finally {
    button.disabled = false;
  }
});
This uses the ONNX model and `q4f16` shown in the Transformers.js 4.3.0 release. Generation passes the input to a pipeline in the page.

Serve the page and check the result and first download

Open a terminal in the folder containing the files and run one of the commands below. Visit `http://127.0.0.1:8000` and click Generate once the status begins with `Ready:`. The first run takes time to download the model; once it is ready, a short result appears for the entered prompt.

Both commands bind to the local computer only. Choose the first if you use Node.js, or the second if Python is installed. Use HTTPS when publishing a public website.

Start a server — choose Node.js or Python
# Option A: Node.js
npx serve --listen tcp://127.0.0.1:8000 .

# Option B: Python (choose one server, not both)
python3 -m http.server 8000 --bind 127.0.0.1
Either command serves the page at `http://127.0.0.1:8000`.

If the model fails to load, check downloads and GPU errors.

If there is no WebGPU adapter, use `q4` for CPU/WASM. For Safari, confirm that it is version 26 or later, the range added in Transformers.js 4.3.0; check WebGPU support separately for other browser versions. A model may still fail to load because of memory limits even when an adapter is returned.

If the model files fail to load, inspect the browser developer tools' Network panel for blocked or failed CDN or Hub requests. For a GPU device error, close other tabs and apps, retry with CPU6 `q4`, or choose a smaller supported model. For errors after input, record the first console error and the model ID, dtype, and device combination to help reproduce the issue.

Return browser extraction results as JSON

JSON can be more convenient than prose when passing an answer to another program. For example, you can extract only sentiment and topic from product feedback. The experimental structured-output feature in 4.3.0 uses JSON Schema to constrain fields and values. It currently generates one sequence at a time.

This feature needs a separate package. Use the following code in a web project with a module bundler such as Vite, rather than pasting it into the earlier CDN HTML. Install the dependencies and add the code to a JavaScript file. Parse the output with `JSON.parse()` and check required fields before using it.

Install dependencies in a bundler project
npm install @huggingface/transformers@4.3.0 @huggingface/transformers-structured-output
Unlike the static HTML example, this code needs a development environment that resolves npm packages.
Constrain sentiment and topic with StructuredOutputProcessor
import { pipeline } from '@huggingface/transformers';
import { StructuredOutputProcessor } from '@huggingface/transformers-structured-output';

const generator = await pipeline('text-generation', 'onnx-community/LFM2.5-350M-ONNX', {
  dtype: 'q4f16', device: 'webgpu',
});

const processor = new StructuredOutputProcessor(generator.tokenizer, {
  type: 'json_schema',
  json_schema: {
    type: 'object',
    properties: {
      sentiment: { enum: ['positive', 'negative', 'neutral'] },
      topic: { enum: ['price', 'quality', 'delivery', 'other'] },
    },
    required: ['sentiment', 'topic'],
    additionalProperties: false,
  },
});

const result = await generator(
  [{ role: 'user', content: 'Classify this feedback: Shipping was fast, but the product is expensive.' }],
  { max_new_tokens: 48, do_sample: false, logits_processor: [processor] },
);
const jsonText = result[0].generated_text.at(-1).content;
console.log(JSON.parse(jsonText));
Run this example in a WebGPU-capable browser.

Terminology notes

  1. Inference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.

    Back to the text
  2. Pipeline — A sequence of processing stages from input to output. Different models or tools may be used at each stage.

    Back to the text
  3. WebGPU — A web-standard API for graphics and general-purpose GPU computation in browsers. Supported features can vary by browser and device.

    Back to the text
  4. System RAM — System memory that temporarily holds data while programs run.

    Back to the text
  5. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  6. CPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.

    Back to the text