Executable programs and extensions
Building Fully Offline Local AI: Air-Gap Installation and Update Procedures
Planning an update path is more important than unplugging the Internet cable.
A fully offline environment doesn't just end with copying the model. You need to prepare the runtime1, dependencies, hashes, licenses, updates and reverts as a bundle.
Include files other than models in the import bundle.
List checkpoints2, tokenizers, runtimes, GPU3 drivers, packages appropriate for your operating system, configuration files, licenses and sources. If any files are automatically downloaded from the Internet during installation, it will stop offline.
Record the hash and version of each file and copy them from the online staging device to a new storage medium. If you go through malware scanning and integrity verification before bringing it inside, you can later track which files were executed.

In an internal environment, we even test reinstallation without a network.
Start with a clean slate to avoid relying on caches that are already installed. Check model loading, first request, automatic startup after reboot, log storage, and failure recovery without external communication.
If you are running a document RAG4 as well, embedding models and indexing tools are also included in the import bundle. Turn off external telemetry and update checks, and make sure internal paths aren't displayed on the user's screen when they fail.

Updates install side by side rather than replace
No new models or runtimes will be overwritten on top of the existing environment. Install the new version in a separate path, pass the same evaluation questions and performance tests, and then just switch the service pointer.
If something goes wrong, you should be able to revert to a previous version. In addition to checkpoints, settings, prompts, index versions, and driver combinations must also be preserved to restore the same state.

Ensure operational records are complete internally
It leaves behind who imported which version, when it was verified, and what data was accessed. Set retention rules appropriate for the level of confidentiality, such as leaving only the request identifier and error level in the log instead of the entire original text.
Just because it's offline doesn't mean it's safe. Because USB import, administrator account, and local tool execution privileges can become new attack vectors, we operate both least privilege and import approval procedures.
Terminology notes
Runtime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.
Back to the textCheckpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.
Back to the textGPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.
Back to the textRetrieval-Augmented Generation — A method that retrieves material relevant to a query, adds it to the model input, and generates an answer. Search scope and source quality depend on the implementation.
Back to the text