Executable programs and extensions

Install Hermes Agent and Connect a Local Model: From Your First Task to Permission Checks

Installing Hermes Agent is different from downloading a Hermes model.

Hermes Agent connects tool calls1, while model weights such as Hermes 4 are prepared separately. First verify the model connection and tools with a short reply and a read-only file listing, then check the endpoint, memory, MCP2, and permissions for external data transfer.

First decide whether you need a model or an agent

Hermes Agent and Hermes models are separate components. To have it read a code folder, prepare the Agent program and the model it will use separately; first identify which one you need to install.

A Hermes model is a set of weights that generates the next tokens3 from an input. Hermes Agent is a program that sends instructions and conversation history to a model and manages tool calls. A model can answer questions on its own, but that does not give it access to local files or the shell. Installing the Agent does not automatically install the model you plan to use.

To fix an error in a project folder, first decide whether you only need a suggested fix or want the program to read and edit files too. That choice changes which components you install, where data goes, and what permissions you grant.

A desk with a blank computer, tool drawer, memory notebook, and geometric task cards in sequence
Hermes Agent connects the model, tools, and memory in one work loop.

After installation, check a reply and a real tool action separately

Suppose you want to find the cause of an error in a project folder. First check that the Agent receives a short reply from the selected model provider. This step tests only the basic connection between the model and endpoint.

Next, ask it to list files in a temporary work folder where only reading is allowed. If the correct list comes back, the Agent called a tool and received its result. This distinguishes ‘it started and answered’ from ‘this task path works.’

For the first test, avoid important repositories and folders containing credentials, and compare the returned list yourself. If it fails, identify whether the endpoint, model response format, or tool configuration stopped the request before granting edit permissions.

Check the current official installation path and begin with a basic task

To install the CLI4 on macOS, Linux, or WSL2, run the command on the official installation page. Windows uses a separate PowerShell command, and the macOS desktop package supports Apple Silicon only. After installation, reopen the terminal and run hermes. Choose a model provider with the hermes model command and configure tools with hermes tools. First check that a short question receives a reply; connect messaging or MCP only when your task needs it.

Install the Hermes CLI on macOS, Linux, or WSL2
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
This installer uses the official domain; Windows uses a separate PowerShell command.

A local model and an OpenAI-compatible endpoint are different connection choices

If you want Hermes Desktop to manage the model on your computer, choose Local Models. The official guide says it is available in Canary builds; other desktop builds may require the --local launch option.

The managed path prepares the llama.cpp runtime5, downloads a model, and checks whether it fits available memory. If you already run an Ollama, llama.cpp, or MLX6 server, you can connect it by entering the server address and model ID in a custom endpoint.

‘OpenAI-compatible’ mainly describes the request format. Separately check whether the model can produce tool calls, the server preserves them, and Hermes executes them. Enter the server’s actual address and API7 path, check that the model list and a short chat work, then call a read-only tool and compare its result.

Consider memory storage separately from the context sent to the model

Hermes uses MEMORY.md and USER.md under ~/.hermes/memories/ for built-in persistent memory. Their contents enter the system prompt at the start of a session, and each file has a character limit.

These files do not store every conversation indefinitely or provide automatic vector search. To find something from an earlier session, check what the separate session-search feature stores and how it searches. Editing a memory file may not change the input of a session that has already started.

Do not let multiple Agent processes edit the same Hermes home memory at the same time. An external memory provider may receive data or model requests, so check its storage location and retention terms separately.

An agent-memory desk separating selected record cards from a closed archive
Curated memory and searchable history are different storage functions.

Review permissions as you add tools, skills, and MCP

Hermes Agent can connect to a terminal, file operations, a browser, skills, and MCP servers. A tool appearing in the list does not mean its access is limited.

Check the commands and package sources used by an MCP server, and leave unnecessary write or delete functions disabled. A skill can include code, commands, or external connections, so inspect its source and behavior. execute_code limits and dangerous-command checks do not replace operating-system isolation.

For code review, start with repository reading and require approval for edits after showing the target and scope. Even when the model runs locally, the Agent has separate file and network permissions.

The word ‘local’ alone does not determine data flow or safety

Even when the Agent runs on your computer, choosing a remote model sends prompts to its provider. A local model can still use other servers for search, a browser, MCP, or messaging. Check where each feature sends prompts, files, and tool results.

Keep API keys out of configuration files and command history, and grant only the permissions required. If another device needs to reach the model server, check authentication and network access; do not expose it directly to the internet.

Hermes terminal approvals are not the same as operating-system isolation. Start by reading a small part of the repository, and review the result before allowing edits or external messages.

A desk where an agent laptop connects to a separate local inference server
Check the agent location and model endpoint location separately.

Terminology notes

  1. Tool call — A structured request from a model for an external function such as reading a file, searching, or running a command. The agent runtime and its permission settings decide whether the request is actually executed.

    Back to the text
  2. MCP — A protocol for connecting clients to tools and data sources a model may use. Connectivity and permission to execute a particular tool must still be configured separately.

    Back to the text
  3. Token — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.

    Back to the text
  4. CLI — Short for Command-Line Interface: operating a program by entering commands in a terminal.

    Back to the text
  5. Runtime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.

    Back to the text
  6. MLX — A machine-learning framework developed by Apple. On Apple silicon it uses unified memory and Metal; separate Linux backends are also available. Model and feature support depends on the MLX-based tool.

    Back to the text
  7. API — A defined interface that lets other code call a program’s functions. The term API alone does not imply sending data to an external server.

    Back to the text