new model · 2026.10.05
Run Xing4.0 Locally: Install the Coding Model and Plan GPU Memory
Before sending a coding task, choose a runtime and GPUs for the full 29B model.
Xing4.0-29B-A4B is a MoE1 agent model with about 4B parameters active per token2 out of 29B total. Its publisher records the BF163, FP84, and GGUF5 releases on 2026-09-17 and provides Apache-2.0 weights and coding/tool-use paths. The 4B active figure does not mean only 4B weights must be stored or loaded. The official vLLM example uses tensor parallelism6 of two and a 256K context, while the repository says upstream runtime7 support PRs are still under review.
A 29B model selects about 4B of experts per token
Xing4.0-29B-A4B is an open-weight agent model designed to read code, propose edits, and continue through tool calls8. Its full model has about 29B parameters; the MoE architecture activates about 4B when computing each token. In an agent framework, the model receives code and error logs and generates a next action or code response, while a separate runner executes the tool call. Publisher’s official model card
We will follow one task: inspect three files in a repository, run a test to investigate a bug, then explain the changes. The model is described as supporting tool calling and a 256K context, but the workflow still needs a code runner, repository permissions, and a supported runtime. Benchmark scores alone do not show that it will edit, test, and revert files reliably.

Do not calculate VRAM from the 4B active figure
Active parameters describe the experts selected for each token’s computation; they do not remove the other experts from memory. A simple BF16 estimate for all weights is 29B × 2 bytes ≈ 58 GB. This is an approximate amount of weight data, not a measured checkpoint9 size or runtime-memory requirement. KV cache10, the runtime, and working headroom add more.
Start memory planning with precision and the actual release file. The official repository records BF16, FP8, and GGUF releases on September 17, 2026. Their file sizes differ, but the existence of a GGUF does not establish that it fits a particular GPU11. Check whether your selected runtime supports both the architecture and that quantized file. Official release record
| Figure | Meaning | Use in hardware planning |
|---|---|---|
| 29B total | Overall checkpoint expert and weight scale | Check file and memory requirements |
| About 4B active | Experts selected for token computation | Use to understand compute per token |
| About 58 GB BF16 arithmetic estimate | Estimated weight data from 29B × 2 bytes | Do not treat as minimum runtime memory |
Separate total weights from per-token computation when planning memory.
29B total
- Meaning
- Overall checkpoint expert and weight scale
- Use in hardware planning
- Check file and memory requirements
About 4B active
- Meaning
- Experts selected for token computation
- Use in hardware planning
- Use to understand compute per token
About 58 GB BF16 arithmetic estimate
- Meaning
- Estimated weight data from 29B × 2 bytes
- Use in hardware planning
- Do not treat as minimum runtime memory

Coding-agent scores depend on the harness and run conditions
The model card reports 75.0 on SWE-bench Verified. The publisher specifies `SWE-agent`, temperature 1.0, top_p 0.95, repetition penalty 1.05, and a 210K context. This score evaluates an agent solving repository tasks, so it is different from a single code-completion response or local tok/s. Xing4.0 official evaluation notes
The Terminal-Bench 2.1 score is 57.5. The card reports a three-run average using `terminus-2`, temperature 0.8, top_p 0.95, repetition penalty 1.05, up to 64K output tokens, and a 24-hour limit per task. This evaluation includes agent tools and a long timeout; it does not show how quickly the model writes one line. On your repository, track edit success, test results, and permitted tools separately.
| Evaluation | Published score | Key conditions |
|---|---|---|
| SWE-bench Verified | 75.0 | SWE-agent, 210K context, temp 1.0; see card for remaining settings |
| Terminal-Bench 2.1 | 57.5 | terminus-2, up to 64K output, 24 hours per task, three-run average |
The published scores come from different coding-agent evaluations.
SWE-bench Verified
- Published score
- 75.0
- Key conditions
- SWE-agent, 210K context, temp 1.0; see card for remaining settings
Terminal-Bench 2.1
- Published score
- 57.5
- Key conditions
- terminus-2, up to 64K output, 24 hours per task, three-run average
Check the pre-merge runtime paths
The publisher’s repository documents Transformers and vLLM, SGLang, and KTransformers configurations. As checked on 2026-10-05, it also says upstream support pull requests for vLLM, SGLang, llama.cpp, and other runtimes are not yet merged. Do not assume the default latest package already includes Xing support; use the PR branch or prebuilt container linked by the official repository. Official run paths and PR status
Keep code-execution permissions separate from the model. Start with read-only access, add test execution next, and allow file edits only when you can inspect the diff. Even when an agent framework connects a repository and terminal, first check whether model outputs execute commands automatically and whether approval is required. Agent scores do not measure the safety of your tool permissions or success on your repository.
python -m venv .venv
source .venv/bin/activate
pip install -U torch transformers accelerate huggingface_hub
hf download XingChen-AGI/Xing4.0-29B-A4BSend a short coding question with the official Transformers example
The official repository’s Transformers example loads the model with `trust_remote_code=True`. This uses model code supplied by the repository, so review the official model files before running it. The example fixes the prompt and caps output at 1,024 tokens as a short check of loading and response format. For repository work, adjust the output limit to the explanation and code changes needed. Xing official Transformers quickstart
On success, the terminal prints a short explanation of the Python error. For an unsupported-architecture error, check Transformers compatibility. For an out-of-memory error, account for the roughly 58 GB arithmetic size of BF16 weights as well as context and runtime memory. Lowering `max_new_tokens` limits generation length but may not fix an error that occurs while loading the weights. Official architecture and inference example
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "XingChen-AGI/Xing4.0-29B-A4B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id, trust_remote_code=True, device_map="auto", dtype=torch.bfloat16
)
messages = [{"role": "user", "content": "Explain this Python error and suggest a minimal fix: KeyError: 'id'"}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs, do_sample=True, top_p=0.95, temperature=0.8,
repetition_penalty=1.05, max_new_tokens=1024
)
print(tokenizer.decode(
output[0][inputs.input_ids.shape[-1]:],
skip_special_tokens=False, spaces_between_special_tokens=False
))The vLLM recipe uses two GPUs and limits concurrent requests
The publisher’s vLLM example sets `--tensor-parallel-size 2`, `--gpu-memory-utilization 0.90`, `--max-model-len 262144`, and `--max-num-seqs 4`. This is a two-GPU sharded serving path with up to four sequences. It does not mean every prompt should use 256K; longer context also uses more KV-cache memory. Decide on prompt length and concurrency, then check whether that configuration fits.
The publisher’s KTransformers example is a heterogeneous path that uses CPU12 and GPU together. It sets the number of experts placed on the GPU and CPU inference13 threads. Thus ‘runs with one GPU’ may mean the selected runtime assigns some work to the CPU, not that all weights fit in GPU memory. The official material does not give a precise minimum RAM14 or inference speed, which must be measured on the chosen setup.
When the server starts, the vLLM API15 responds at `http://localhost:8000/v1`. If startup reports an architecture or parser error, check that you are using the publisher’s prebuilt image or a PR branch linked by the repository rather than a generic vLLM build. For an out-of-memory error, first reduce context and concurrent sequences, then check that both GPUs are visible. Runtime after such changes remains unmeasured until you test it directly. Official vLLM settings and support paths
| Path | Published setup | Still to verify |
|---|---|---|
| Transformers | device_map=auto, BF16 example | Actual placement memory and supported hardware |
| vLLM | TP=2, 256K maximum context, up to 4 sequences | Upstream PR not merged; use official PR branch/image |
| KTransformers | CPU/GPU heterogeneous example | System RAM and runtime are not published |
Compare the execution assumptions in the official serving paths.
Transformers
- Published setup
- device_map=auto, BF16 example
- Still to verify
- Actual placement memory and supported hardware
vLLM
- Published setup
- TP=2, 256K maximum context, up to 4 sequences
- Still to verify
- Upstream PR not merged; use official PR branch/image
KTransformers
- Published setup
- CPU/GPU heterogeneous example
- Still to verify
- System RAM and runtime are not published

Judge fit with a small repository task
After the Python `KeyError: 'id'` explanation appears, use a small repository bug as the next read-only task. Send the reproduction command and relevant files, then ask which test to run. Before letting the agent edit files, check that its file path, cause, and verification steps agree. Only then allow edits in a limited directory and have a person review the diff and test results.
For hardware selection, account for storage for the checkpoint, memory for weights and input, and headroom for the apps you normally use. The official TP=2 path requires two GPUs and its runtime setup. A KTransformers CPU/GPU path uses other memory resources when GPU memory is limited, but completion time may change. If the task does not require this exact model, compare it with a smaller coding model on the same repository issue before spending on hardware.
Terminology notes
MoE — A model architecture that selects some of several expert subnetworks for each input. Total parameters can differ from the number activated for one token.
Back to the textToken — A unit into which a model divides input or output for processing.
Back to the textBF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.
Back to the textFP8 — A family of 8-bit floating-point formats. Specific formats and support vary by hardware and software.
Back to the textGGUF — A file format for model data, widely used by llama.cpp-based tools.
Back to the textTensor parallelism — Splitting computations and weights within model layers across devices. Communication is required, so speedups depend on the interconnect and implementation.
Back to the textRuntime — The software environment that provides facilities needed while a program runs. In local AI, it can also refer to a model execution engine.
Back to the textTool call — A structured request from a model for an external function such as reading a file, searching, or running a command. The agent runtime and its permission settings decide whether the request is actually executed.
Back to the textCheckpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.
Back to the textKV cache — Memory that stores attention keys and values from earlier tokens for reuse during later token generation. Its size depends on context length and batch size.
Back to the textGPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.
Back to the textCPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.
Back to the textInference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.
Back to the textSystem RAM — System memory that temporarily holds data while programs run.
Back to the textAPI — A defined interface that lets other code call a program’s functions.
Back to the text
Read next
read first
A computer for local coding AI: autocomplete and agents need different tests
read first
Hardware for Local AI Agents: Choose by Workload, Context, and Uptime
Executable programs and extensions
Running local LLMs across two or more GPUs
Response speed and acceleration
MTP: how multi-token prediction works and when to use it
