Clear, practical technology insights BSOD Code Lookup · Windows Error Code Lookup · Wi-Fi Troubleshooting · PC Troubleshooting Checklist

How to Run Qwen3.5-397B on One GPU with llama.cpp

Build llama.cpp with CUDA, download the six-part Qwen3.5 GGUF, start a one-GPU server with RAM offloading, and access it through SSH.

Table of Contents

You can run the Qwen3.5-397B-A17B GGUF with llama.cpp on one NVIDIA GPU by keeping part of the model in system RAM. The example here uses one H200-class GPU and substantial host memory on a Linux machine. If you rent that machine, inference runs on the remote host; an SSH tunnel makes its web UI and API available at 127.0.0.1 on your laptop. This is self-hosted inference, not execution on the laptop’s hardware.

Check memory, disk, and software first

The Unsloth MXFP4_MOE GGUF is roughly 237 GB on disk and comes in six shards. Plan for more free storage than the download size and enough combined GPU and system memory for weights, runtime buffers, and the context cache. A single H200 with around 141 GB of VRAM and a host with roughly 240 GB of RAM is the example configuration, not a guaranteed minimum: speed and memory use vary with offloading and context. If your machine is smaller, consider a smaller Qwen3.5 variant instead of assuming SSD swap will make the 397B model practical.

Use a Linux host with a working NVIDIA driver, CUDA development tools, CMake, a C++ compiler, Git, and Python. Confirm CUDA availability with nvidia-smi and nvcc --version before building. Review the Unsloth Qwen3.5 guide and the llama.cpp build instructions if your driver or build environment differs.

1. Connect to the GPU host

On a local workstation with adequate hardware, work directly in its terminal. For a rented Linux VM, choose a single-GPU instance with enough RAM and disk, add your SSH public key, and note its actual user name and IP address. The screenshots below show one provider’s setup; its prices and availability can change.

How to run Qwen 3.5 locally on a single GPU Picture 1

How to run Qwen 3.5 locally on a single GPU Picture 2

From your computer, connect with an SSH tunnel. Replace the placeholders with the credentials for your machine:

ssh -L 8080:127.0.0.1:8080 USER@SERVER_IP

Verify the server’s SSH host key before accepting a first connection. Keep this session open while you use the forwarded UI.

How to run Qwen 3.5 locally on a single GPU Picture 3

On the VM, check the GPU:

nvidia-smi
nvcc --version

If nvcc is unavailable, install a compatible CUDA Toolkit for that host before compiling the CUDA backend.

How to run Qwen 3.5 locally on a single GPU Picture 4

2. Build llama.cpp with CUDA

On an Ubuntu-like host, install basic build and download tools, clone the current repository, and compile llama-server. Run each command separately; the source instructions had two package commands accidentally joined together.

sudo apt update
sudo apt install -y git build-essential cmake curl libcurl4-openssl-dev python3-venv
git clone https://github.com/ggml-org/llama.cpp
cmake -S llama.cpp -B llama.cpp/build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build llama.cpp/build --config Release -j --target llama-server

CMake can select the attached GPU’s architecture; specify CMAKE_CUDA_ARCHITECTURES only when your build requires it. The binary stays at llama.cpp/build/bin/llama-server, so there is no need to copy files into the source tree.

How to run Qwen 3.5 locally on a single GPU Picture 5

How to run Qwen 3.5 locally on a single GPU Picture 6

3. Download every shard of the chosen GGUF

Install the Hugging Face CLI in a Python virtual environment on the GPU host, then download only the MXFP4_MOE files into a predictable directory:

python3 -m venv ~/hf-tools
~/hf-tools/bin/python -m pip install -U huggingface_hub
~/hf-tools/bin/hf download unsloth/Qwen3.5-397B-A17B-GGUF --local-dir models/Qwen3.5 --include "MXFP4_MOE/*"

Check that all six .gguf shards are present in models/Qwen3.5/MXFP4_MOE/. The first shard is the model path you pass to the server; llama.cpp reads the other shards from the same directory. The download consumes storage on the host, not GPU VRAM.

How to run Qwen 3.5 locally on a single GPU Picture 7

4. Start the server and test the tunnel

Run this on the GPU host from the directory containing llama.cpp and models. Bind to loopback so the unauthenticated server is reachable through SSH rather than from the public internet:

./llama.cpp/build/bin/llama-server --model models/Qwen3.5/MXFP4_MOE/Qwen3.5-397B-A17B-MXFP4_MOE-00001-of-00006.gguf --alias Qwen3.5 --host 127.0.0.1 --port 8080 --fit on --ctx-size 8192 --jinja

--fit on asks llama.cpp to adjust unspecified parameters to available device memory; it does not create RAM or make an undersized host fast. A larger context consumes more memory. If loading fails, check the downloaded shards, current build support, free RAM/VRAM, and server log before changing offload or context settings. If you later need image input, download the matching mmproj file and pass it with --mmproj; the text-only test below does not need it.

How to run Qwen 3.5 locally on a single GPU Picture 8

How to run Qwen 3.5 locally on a single GPU Picture 9

Keep the server and SSH tunnel running. On your laptop, open http://127.0.0.1:8080 for the web UI or test the API in another terminal:

curl -s http://127.0.0.1:8080/v1/models

Seeing the Qwen3.5 alias confirms that the local tunnel reaches the server. To test generation, send a small request:

curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"Qwen3.5","messages":[{"role":"user","content":"Write one sentence about AI agents."}]}'

Wait for a response before assuming the model is ready; a listed model alone does not prove that generation succeeds. Other clients that support this style of chat endpoint can point to http://127.0.0.1:8080/v1.

How to run Qwen 3.5 locally on a single GPU Picture 10

How to run Qwen 3.5 locally on a single GPU Picture 11

5. Try the web UI with a small application prompt

The built-in UI lets you test the model without client code. The following screenshots show a prompt for a Python terminal stock screener and its generated output. Treat this as a code-generation exercise, not as validated financial analysis. Ask the model to define its inputs, data source, scoring rule, and output file; review the generated code and test it on sample data before installing packages or running it.

How to run Qwen 3.5 locally on a single GPU Picture 12

How to run Qwen 3.5 locally on a single GPU Picture 13

The example uses Python libraries such as Rich for a terminal interface and yfinance for public market data. Data fields and provider availability can change. Inspect network calls, errors, missing data, and CSV output, and never assume a generated ranking is suitable for an investment decision.

How to run Qwen 3.5 locally on a single GPU Picture 14

How to run Qwen 3.5 locally on a single GPU Picture 15

For a lighter local setup, see the smaller Qwen 3 Ollama guide. If you prefer a desktop interface for compatible GGUF models, read the LM Studio setup guide.

Discussion

Reader Comments 0

Sign in with email or Google to join the discussion.