new model · 2026.10.05

Run Xing4.0 Locally: Install the Coding Model and Plan GPU Memory

Before sending a coding task, choose a runtime and GPUs for the full 29B model.

Xing4.0-29B-A4B is a MoE1 agent model with about 4B parameters active per token2 out of 29B total. Its publisher records the BF163, FP84, and GGUF5 releases on 2026-09-17 and provides Apache-2.0 weights and coding/tool-use paths. The 4B active figure does not mean only 4B weights must be stored or loaded. The official vLLM example uses tensor parallelism6 of two and a 256K context, while the repository says upstream runtime7 support PRs are still under review.

Requirements and key details
  • The exact model ID is `XingChen-AGI/Xing4.0-29B-A4B`, under Apache-2.0.
  • The 29B total and 4B active parameter counts describe different things for memory planning.
  • The official vLLM configuration example uses TP=2, 256K context, and up to four sequences.
  • SWE-bench and Terminal-Bench scores measure agent tasks, not tok/s.

A 29B model selects about 4B of experts per token

Xing4.0-29B-A4B is an open-weight agent model designed to read code, propose edits, and continue through tool calls8. Its full model has about 29B parameters; the MoE architecture activates about 4B when computing each token. In an agent framework, the model receives code and error logs and generates a next action or code response, while a separate runner executes the tool call. Publisher’s official model card

We will follow one task: inspect three files in a repository, run a test to investigate a bug, then explain the changes. The model is described as supporting tool calling and a 256K context, but the workflow still needs a code runner, repository permissions, and a supported runtime. Benchmark scores alone do not show that it will edit, test, and revert files reliably.

A code editor, terminal window, and notepad sit open on a watercolor-painted coding desk.
The agent uses both code and run output to propose a next step.

Do not calculate VRAM from the 4B active figure

Active parameters describe the experts selected for each token’s computation; they do not remove the other experts from memory. A simple BF16 estimate for all weights is 29B × 2 bytes ≈ 58 GB. This is an approximate amount of weight data, not a measured checkpoint9 size or runtime-memory requirement. KV cache10, the runtime, and working headroom add more.

Start memory planning with precision and the actual release file. The official repository records BF16, FP8, and GGUF releases on September 17, 2026. Their file sizes differ, but the existence of a GGUF does not establish that it fits a particular GPU11. Check whether your selected runtime supports both the architecture and that quantized file. Official release record

Separate total weights from per-token computation when planning memory.
FigureMeaningUse in hardware planning
29B totalOverall checkpoint expert and weight scaleCheck file and memory requirements
About 4B activeExperts selected for token computationUse to understand compute per token
About 58 GB BF16 arithmetic estimateEstimated weight data from 29B × 2 bytesDo not treat as minimum runtime memory

Separate total weights from per-token computation when planning memory.

29B total

Meaning
Overall checkpoint expert and weight scale
Use in hardware planning
Check file and memory requirements

About 4B active

Meaning
Experts selected for token computation
Use in hardware planning
Use to understand compute per token

About 58 GB BF16 arithmetic estimate

Meaning
Estimated weight data from 29B × 2 bytes
Use in hardware planning
Do not treat as minimum runtime memory
An open desktop case reveals a large GPU beside a laptop showing model settings, in a watercolor scene.
Plan for 4B active computation separately from 29B total weights.

Coding-agent scores depend on the harness and run conditions

The model card reports 75.0 on SWE-bench Verified. The publisher specifies `SWE-agent`, temperature 1.0, top_p 0.95, repetition penalty 1.05, and a 210K context. This score evaluates an agent solving repository tasks, so it is different from a single code-completion response or local tok/s. Xing4.0 official evaluation notes

The Terminal-Bench 2.1 score is 57.5. The card reports a three-run average using `terminus-2`, temperature 0.8, top_p 0.95, repetition penalty 1.05, up to 64K output tokens, and a 24-hour limit per task. This evaluation includes agent tools and a long timeout; it does not show how quickly the model writes one line. On your repository, track edit success, test results, and permitted tools separately.

The published scores come from different coding-agent evaluations.
EvaluationPublished scoreKey conditions
SWE-bench Verified75.0SWE-agent, 210K context, temp 1.0; see card for remaining settings
Terminal-Bench 2.157.5terminus-2, up to 64K output, 24 hours per task, three-run average

The published scores come from different coding-agent evaluations.

SWE-bench Verified

Published score
75.0
Key conditions
SWE-agent, 210K context, temp 1.0; see card for remaining settings

Terminal-Bench 2.1

Published score
57.5
Key conditions
terminus-2, up to 64K output, 24 hours per task, three-run average

Check the pre-merge runtime paths

The publisher’s repository documents Transformers and vLLM, SGLang, and KTransformers configurations. As checked on 2026-10-05, it also says upstream support pull requests for vLLM, SGLang, llama.cpp, and other runtimes are not yet merged. Do not assume the default latest package already includes Xing support; use the PR branch or prebuilt container linked by the official repository. Official run paths and PR status

Keep code-execution permissions separate from the model. Start with read-only access, add test execution next, and allow file edits only when you can inspect the diff. Even when an agent framework connects a repository and terminal, first check whether model outputs execute commands automatically and whether approval is required. Agent scores do not measure the safety of your tool permissions or success on your repository.

Prepare the official Transformers environment
python -m venv .venv
source .venv/bin/activate
pip install -U torch transformers accelerate huggingface_hub
hf download XingChen-AGI/Xing4.0-29B-A4B
These commands prepare the Transformers path documented by the repository. Use the Python example in the next section to load the model. Installation does not establish that a 29B BF16 model fits; inspect the checkpoint files and available memory first.

Send a short coding question with the official Transformers example

The official repository’s Transformers example loads the model with `trust_remote_code=True`. This uses model code supplied by the repository, so review the official model files before running it. The example fixes the prompt and caps output at 1,024 tokens as a short check of loading and response format. For repository work, adjust the output limit to the explanation and code changes needed. Xing official Transformers quickstart

On success, the terminal prints a short explanation of the Python error. For an unsupported-architecture error, check Transformers compatibility. For an out-of-memory error, account for the roughly 58 GB arithmetic size of BF16 weights as well as context and runtime memory. Lowering `max_new_tokens` limits generation length but may not fix an error that occurs while loading the weights. Official architecture and inference example

Load the model and check a capped response
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "XingChen-AGI/Xing4.0-29B-A4B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, trust_remote_code=True, device_map="auto", dtype=torch.bfloat16
)

messages = [{"role": "user", "content": "Explain this Python error and suggest a minimal fix: KeyError: 'id'"}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.inference_mode():
    output = model.generate(
        **inputs, do_sample=True, top_p=0.95, temperature=0.8,
        repetition_penalty=1.05, max_new_tokens=1024
    )
print(tokenizer.decode(
    output[0][inputs.input_ids.shape[-1]:],
    skip_special_tokens=False, spaces_between_special_tokens=False
))
This requires a supported hardware and Transformers combination. On success, generated text appears in the terminal. If the architecture is reported as unsupported, first check that the installed Transformers release supports Xing and that the checkpoint path and dtype are correct.

The vLLM recipe uses two GPUs and limits concurrent requests

The publisher’s vLLM example sets `--tensor-parallel-size 2`, `--gpu-memory-utilization 0.90`, `--max-model-len 262144`, and `--max-num-seqs 4`. This is a two-GPU sharded serving path with up to four sequences. It does not mean every prompt should use 256K; longer context also uses more KV-cache memory. Decide on prompt length and concurrency, then check whether that configuration fits.

The publisher’s KTransformers example is a heterogeneous path that uses CPU12 and GPU together. It sets the number of experts placed on the GPU and CPU inference13 threads. Thus ‘runs with one GPU’ may mean the selected runtime assigns some work to the CPU, not that all weights fit in GPU memory. The official material does not give a precise minimum RAM14 or inference speed, which must be measured on the chosen setup.

When the server starts, the vLLM API15 responds at `http://localhost:8000/v1`. If startup reports an architecture or parser error, check that you are using the publisher’s prebuilt image or a PR branch linked by the repository rather than a generic vLLM build. For an out-of-memory error, first reduce context and concurrent sequences, then check that both GPUs are visible. Runtime after such changes remains unmeasured until you test it directly. Official vLLM settings and support paths

Compare the execution assumptions in the official serving paths.
PathPublished setupStill to verify
Transformersdevice_map=auto, BF16 exampleActual placement memory and supported hardware
vLLMTP=2, 256K maximum context, up to 4 sequencesUpstream PR not merged; use official PR branch/image
KTransformersCPU/GPU heterogeneous exampleSystem RAM and runtime are not published

Compare the execution assumptions in the official serving paths.

Transformers

Published setup
device_map=auto, BF16 example
Still to verify
Actual placement memory and supported hardware

vLLM

Published setup
TP=2, 256K maximum context, up to 4 sequences
Still to verify
Upstream PR not merged; use official PR branch/image

KTransformers

Published setup
CPU/GPU heterogeneous example
Still to verify
System RAM and runtime are not published
Code-review desk with a laptop and checking notes
A person reviews the diff and tests after the model responds.

Judge fit with a small repository task

After the Python `KeyError: 'id'` explanation appears, use a small repository bug as the next read-only task. Send the reproduction command and relevant files, then ask which test to run. Before letting the agent edit files, check that its file path, cause, and verification steps agree. Only then allow edits in a limited directory and have a person review the diff and test results.

For hardware selection, account for storage for the checkpoint, memory for weights and input, and headroom for the apps you normally use. The official TP=2 path requires two GPUs and its runtime setup. A KTransformers CPU/GPU path uses other memory resources when GPU memory is limited, but completion time may change. If the task does not require this exact model, compare it with a smaller coding model on the same repository issue before spending on hardware.

Terminology notes

  1. MoE — A model architecture that selects some of several expert subnetworks for each input. Total parameters can differ from the number activated for one token.

    Back to the text
  2. Token — A unit into which a model divides input or output for processing.

    Back to the text
  3. BF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.

    Back to the text
  4. FP8 — A family of 8-bit floating-point formats. Specific formats and support vary by hardware and software.

    Back to the text
  5. GGUF — A file format for model data, widely used by llama.cpp-based tools.

    Back to the text
  6. Tensor parallelism — Splitting computations and weights within model layers across devices. Communication is required, so speedups depend on the interconnect and implementation.

    Back to the text
  7. Runtime — The software environment that provides facilities needed while a program runs. In local AI, it can also refer to a model execution engine.

    Back to the text
  8. Tool call — A structured request from a model for an external function such as reading a file, searching, or running a command. The agent runtime and its permission settings decide whether the request is actually executed.

    Back to the text
  9. Checkpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.

    Back to the text
  10. KV cache — Memory that stores attention keys and values from earlier tokens for reuse during later token generation. Its size depends on context length and batch size.

    Back to the text
  11. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  12. CPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.

    Back to the text
  13. Inference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.

    Back to the text
  14. System RAM — System memory that temporarily holds data while programs run.

    Back to the text
  15. API — A defined interface that lets other code call a program’s functions.

    Back to the text