Executable programs and extensions
Run a Local LLM in Your Browser with Transformers.js
Build a small web app that answers prompts inside your browser.
Transformers.js runs AI models from JavaScript. Once an ONNX model is downloaded, the browser can generate replies locally. This guide builds a prompt field and a reply button using LFM2.5-350M, the official example model for version 4.3.0.
What changes with browser-local inference?
A typical AI website sends a prompt to a server and receives a reply. Browser-local inference1 moves that work to the user's computer. Create a Transformers.js `pipeline2()` once, then pass it a prompt: the downloaded model generates a reply inside the page.
No separate Python inference server is needed, making this useful for testing small web features. Start with lightweight tasks such as text classification or short replies. Consider a dedicated serving engine for many concurrent users or continuous large-model serving.

Choose a model and device supported by 4.3.0
Transformers.js 4.3.0 was released on September 16, 2026. Its release example uses `q4f16` for WebGPU3 with `onnx-community/LFM2.5-350M-ONNX`. The model repository also has a `q4` file for CPU/WASM, but no `q8` file was listed, so the code below selects between the two available formats based on the device.
When switching to another ONNX model, confirm that the selected dtype file exists in that repository. The file and inference settings must match for the model to load.
Official sources do not specify minimum system RAM4 or GPU5 memory for this model. Do not infer whether it will run from the model file size alone; check the tab's memory use in the browser task manager during the first run.
| Execution path | Setting | When to use |
|---|---|---|
| CPU/WASM | Omit `device`, use `dtype: 'q4'` | Default inference path using the q4 file present in this repository |
| Supported GPU | `device: 'webgpu'`, `dtype: 'q4f16'` | When a WebGPU adapter is available and supports this format |
| q8 | Do not select for this model | No q8 ONNX file is listed in the model repository |
These are the execution formats and browser device paths confirmed in this model repository.
CPU/WASM
- Setting
- Omit `device`, use `dtype: 'q4'`
- When to use
- Default inference path using the q4 file present in this repository
Supported GPU
- Setting
- `device: 'webgpu'`, `dtype: 'q4f16'`
- When to use
- When a WebGPU adapter is available and supports this format
q8
- Setting
- Do not select for this model
- When to use
- No q8 ONNX file is listed in the model repository

Prepare a static HTML page on localhost
You only need two static files to start. Because the page uses an ES module CDN import, serve it with a local web server instead of opening the file directly from a file browser. `localhost` is treated as a secure context for WebGPU. Use HTTPS for a public deployment.
In a new empty folder, create `index.html` and `main.js`, then save the HTML below. This follows the official documented approach of importing the library as a module from jsDelivr. Pinning the version in the URL makes this example request 4.3.0 even if a later CDN tag changes.
<!doctype html>
<html lang="en">
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Browser-local text generation</title>
<label>Prompt <input id="prompt" value="Reply in one short sentence: local AI runs in a browser."></label>
<button id="run" disabled>Generate</button>
<pre id="status">Loading model…</pre>
<pre id="output"></pre>
<script type="module" src="./main.js"></script>
</html>
Load the model and generate once
The following `main.js` checks for WebGPU before loading the model. It selects `q4f16` when a GPU adapter is available, or CPU/WASM `q4` if no adapter is available or the check fails. The Generate button becomes enabled when the model is ready.
On the first run, the library is fetched from a CDN and model files from Hugging Face Hub. This example processes prompts in the browser rather than sending them to a separate inference server. When integrating it into a service, also check whether analytics or logging code transmits input. Load the model once and reuse the same pipeline for each click.
import { pipeline } from 'https://cdn.jsdelivr.net/npm/@huggingface/transformers@4.3.0';
const status = document.querySelector('#status');
const output = document.querySelector('#output');
const promptInput = document.querySelector('#prompt');
const button = document.querySelector('#run');
const modelId = 'onnx-community/LFM2.5-350M-ONNX';
let adapter = null;
try {
adapter = await navigator.gpu?.requestAdapter() ?? null;
} catch (error) {
console.warn('WebGPU unavailable; using CPU/WASM.', error);
}
const device = adapter ? 'webgpu' : undefined;
const dtype = adapter ? 'q4f16' : 'q4';
let generator;
try {
status.textContent = `Loading ${modelId} (${device ?? 'CPU/WASM'}, ${dtype})…`;
generator = await pipeline('text-generation', modelId, { dtype, ...(device ? { device } : {}) });
status.textContent = `Ready: ${device ?? 'CPU/WASM'} / ${dtype}`;
button.disabled = false;
} catch (error) {
status.textContent = `Model load failed: ${error.message}`;
console.error(error);
}
button.addEventListener('click', async () => {
if (!generator) return;
button.disabled = true;
status.textContent = 'Generating…';
try {
const result = await generator(promptInput.value, { max_new_tokens: 48 });
const generated = result[0]?.generated_text;
output.textContent = typeof generated === 'string'
? generated
: Array.isArray(generated)
? generated.at(-1)?.content ?? JSON.stringify(result)
: JSON.stringify(result);
status.textContent = 'Done.';
} catch (error) {
status.textContent = `Generation failed: ${error.message}`;
console.error(error);
} finally {
button.disabled = false;
}
});Serve the page and check the result and first download
Open a terminal in the folder containing the files and run one of the commands below. Visit `http://127.0.0.1:8000` and click Generate once the status begins with `Ready:`. The first run takes time to download the model; once it is ready, a short result appears for the entered prompt.
Both commands bind to the local computer only. Choose the first if you use Node.js, or the second if Python is installed. Use HTTPS when publishing a public website.
# Option A: Node.js
npx serve --listen tcp://127.0.0.1:8000 .
# Option B: Python (choose one server, not both)
python3 -m http.server 8000 --bind 127.0.0.1If the model fails to load, check downloads and GPU errors.
If there is no WebGPU adapter, use `q4` for CPU/WASM. For Safari, confirm that it is version 26 or later, the range added in Transformers.js 4.3.0; check WebGPU support separately for other browser versions. A model may still fail to load because of memory limits even when an adapter is returned.
If the model files fail to load, inspect the browser developer tools' Network panel for blocked or failed CDN or Hub requests. For a GPU device error, close other tabs and apps, retry with CPU6 `q4`, or choose a smaller supported model. For errors after input, record the first console error and the model ID, dtype, and device combination to help reproduce the issue.
Return browser extraction results as JSON
JSON can be more convenient than prose when passing an answer to another program. For example, you can extract only sentiment and topic from product feedback. The experimental structured-output feature in 4.3.0 uses JSON Schema to constrain fields and values. It currently generates one sequence at a time.
This feature needs a separate package. Use the following code in a web project with a module bundler such as Vite, rather than pasting it into the earlier CDN HTML. Install the dependencies and add the code to a JavaScript file. Parse the output with `JSON.parse()` and check required fields before using it.
npm install @huggingface/transformers@4.3.0 @huggingface/transformers-structured-outputimport { pipeline } from '@huggingface/transformers';
import { StructuredOutputProcessor } from '@huggingface/transformers-structured-output';
const generator = await pipeline('text-generation', 'onnx-community/LFM2.5-350M-ONNX', {
dtype: 'q4f16', device: 'webgpu',
});
const processor = new StructuredOutputProcessor(generator.tokenizer, {
type: 'json_schema',
json_schema: {
type: 'object',
properties: {
sentiment: { enum: ['positive', 'negative', 'neutral'] },
topic: { enum: ['price', 'quality', 'delivery', 'other'] },
},
required: ['sentiment', 'topic'],
additionalProperties: false,
},
});
const result = await generator(
[{ role: 'user', content: 'Classify this feedback: Shipping was fast, but the product is expensive.' }],
{ max_new_tokens: 48, do_sample: false, logits_processor: [processor] },
);
const jsonText = result[0].generated_text.at(-1).content;
console.log(JSON.parse(jsonText));Terminology notes
Inference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.
Back to the textPipeline — A sequence of processing stages from input to output. Different models or tools may be used at each stage.
Back to the textWebGPU — A web-standard API for graphics and general-purpose GPU computation in browsers. Supported features can vary by browser and device.
Back to the textSystem RAM — System memory that temporarily holds data while programs run.
Back to the textGPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.
Back to the textCPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.
Back to the text
Read next
read first
What is a local LLM? Setup, PC requirements, models and speed
Executable programs and extensions
llama.cpp: finding which settings are slowing you down
Executable programs and extensions
Backburner: Accelerate Mac LLM Prefill with an iPhone
Executable programs and extensions
When local RAG answer is wrong: Search, chunk, model diagnosis order
