GLM 5.3 (full) EXL3, TP6 and TP4
Serving the full GLM 5.3 model with vLLM across six or four DGX Spark nodes.
- TP6 prefill 1000 tok/s at 8K
- TP6 prose decode 27.7 tok/s (shipped, MTP k=4)
- 360K max context
We put AI on hardware you own, wire it into the software you already run, and automate the work your team keeps redoing by hand. Everything we build has to go into daily use. That is the only test that counts.
The call is free and takes about 20 minutes. If you already know your problem and want to start working, book the paid hour instead.
We run large models on our own NVIDIA DGX Spark cluster and publish what works: launch scripts, overlays, measured numbers, and the caveats. Free to use, on GitHub and Hugging Face.
Serving the full GLM 5.3 model with vLLM across six or four DGX Spark nodes.
A vLLM serving recipe across two DGX Spark boxes over their ConnectX-7 RoCE link. No image rebuild.
The per-rank int4 GPTQ codes used by the GLM 5.3 TP6 and TP4 builds for the dense decode path.
Numbers are single-request measurements on our cluster; each README has the conditions and caveats. See everything on GitHub →
Bring your questions, leave with a plan. A working session with the person who would build the system, not a sales call.
Most companies do not need an AI strategy. They need one system that works, and then another one. Start wherever the pain is.
An AI machine that lives in your office. We deploy NVIDIA class local AI hardware, connect it to your documents and systems, and give your whole team access. Nothing goes to an outside model.
Custom software that connects your business to the big cloud models without handing over the keys. Your IT people can see exactly what leaves the building and what comes back.
The quotes, intake forms, invoices, reports, and follow ups that eat your week. We find the repeat work, automate it, and show you the hours it gives back.
Every AI tool your team signs up for bills by the seat, every month, with no end date. Stack four of them and you are paying rent on software that also keeps a copy of your work.
Buy the machine once instead. Everyone in the building uses it. Your bill does not move when you hire five more people, and the data sits somewhere you can point at.
We are not against cloud AI. Half of what we build runs on it, because for some jobs the frontier models are worth every penny. But four overlapping per seat tools nobody chose on purpose is not a strategy. It is a leak.
What a private AI system replacesIllustrative math for a 40 person company at typical per seat prices. Your numbers will be different. The shape of the problem usually is not.
Most of the AI shops in this market opened in the last two years. Here is what sits behind this one.
This is the AI division of an established Coeur d'Alene firm that has been building, hosting, and supporting business systems for years. Same company, same phone number, same people to yell at.
AI lands on your network and your servers whether anyone planned for it or not. We bring in senior IT and cybersecurity people and vendor partners to handle that side instead of calling it somebody else's problem.
A proof of concept that nobody opens after week two is wasted money. What we build goes into daily operation, with training, documentation, and a person on the phone when it breaks.
The hardware, the configuration, the documentation, the credentials. All yours. Fire us and the system keeps running, which is exactly how we set it up.
Chuck has been teaching business owners how to use AI on his YouTube channel, Chuck the Contractor, since 2023. The audience there is contractors, so the examples are job sites and estimates. The systems underneath are the ones we build here.
More of it is on the channel. If you want the version aimed at your industry instead of contractors, that is what the call is for.
You should be using something real inside the first month. Anyone who wants a 90 day discovery phase before you see software is billing you to learn your business.
A free 20 minute call to see whether there is anything here worth doing. Or the paid hour, if you already know what you want to fix.
We map the workflows, the data, and the risk. You get a written plan: what to build first, what it costs, what it should give back, and which parts stay in the building.
Hardware, integrations, custom software, training. Scoped and priced before we start, and you use a rough version early instead of waiting for a reveal at the end.
Monitoring, updates, backups, model swaps, and the next workflow when you find it. Month to month, at a number you know in advance.
We are in Coeur d'Alene. When the AI runs on a machine in your building, it helps that the people who installed it can be standing next to it in twenty minutes.
Coeur d'Alene, Post Falls, Hayden, Rathdrum, Sandpoint, Spokane, Spokane Valley, Liberty Lake: we work those in person. Everywhere else we work remotely, and we will tell you up front which parts of a job actually need somebody on site. Most of it does not.
This market is moving. The University of Idaho is starting AI degrees in Coeur d'Alene this fall, manufacturers around Spokane are already running machine vision on the floor, and there is a technology summit at the resort in October. The companies that sort their systems out this year are the ones quoting faster than their competition next year.
We answer within one business day. If this is not something we should be doing for you, we will say so on the first call and point you at whoever should.
(208) 648-3977
212 S 11th St Unit #4A
Coeur d'Alene, ID 83814
First name, last name, email, and phone are required. We do not share your information.