The coding model just moved to a single gpu. china telecom's xing4.0-29b-a4b, out september 20, is…
the coding model just moved to a single gpu. china telecom's xing4.0-29b-a4b, out september 20, is a full-stack domestic open-weight coding agent model that fits on one rtx 3090, built for enterprises whose data cannot leave the building.
local-first ai is no longer a compromise.
Context
The Xing4.0-29B-A4B model card on Hugging Face describes 29B total and 4B active parameters, 256K context extensible to 512K, an agent-oriented architecture and training entirely on the Ascend NPU platform with MindSpore. Its GitHub README news line reads 2026-09-17 open-sourced, with an FP8 variant the same day, and lists llama.cpp and SGLang support as pending. A third-party hardware blog of 18 September says the official 4-bit GGUF is 18.72 GiB, that the vendor's llama.cpp guide was tested on an RTX 3090 at 64K context, that it cannot run in stock llama.cpp, Ollama or LM Studio yet and that it found no public tokens-per-second figure on a 3090. The card lists vendor-reported SWE-bench Verified 75.00 and Terminal-Bench 2.1 57.50.
The first-party date is 17 September and not 20 September. Specs and Ascend training are first-party. Fits on one RTX 3090 rests on the third-party blog and was not confirmed in first-party text, and it needs a fork or the vendor's package today. The Apache-2.0 license appears on the third-party blog only and is unconfirmed. In the vendor's own table Qwen3.6-35B-A3B scores higher on SWE-bench Verified at 76.00 while Xing is higher on Terminal-Bench 2.1, and these are vendor-run figures. No first-party data-residency statement was found, so built for enterprises whose data cannot leave the building is the author's framing. That local-first is no longer a compromise is the author's opinion.
Related work
- Earlier note on a larger local model ↗GLM-5.3-Flash, same local-run theme.
Watch next
- The llama.cpp support merge, a first-party hardware guide and independent coding evaluations.
Sources
- Xing4.0-29B-A4B (Hugging Face)huggingface.co
- Xing4.0-29B-A4B README (GitHub)raw.githubusercontent.com
- Can I run Xing 4.0 29B A4B on an RTX 3090 (OpenClaw DC, 18 Sep 2026)openclawdc.com
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 18:02 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →