About the practice

Built for the summer the calculus flipped.

Chels.ai is a sovereign AI infrastructure practice: we design, build, fine-tune, host, and maintain on-premise deployments of flagship open-weight models for organizations that can’t send their data to someone else’s cloud — law firms, clinics, credit unions, and manufacturers across Canada.law firms, health networks, banks, and defense suppliers across the United States.

Why this firm exists

Five weeks in 2026 changed who gets to own frontier intelligence.

This practice wasn’t founded on a trend deck. It was founded on four dated events — and on what they mean for every organization that depends on AI it doesn’t own.

June 11, 2026

Washington cuts off foreign access to top US models

The US government ordered the leading American lab to block foreign access to its most advanced models within 48 hours, with no avenue for appeal. Organizations outside the US that had built workflows on rented frontier intelligence lost it overnight.

June 13, 2026

GLM-5.2 goes open source under MIT

Two days later, Zhipu AI released GLM-5.2 — 744B parameters, 1M-token context — under the fully permissive MIT license, with no regional restrictions. A frontier-adjacent model anyone could download, audit, and run.

July 16, 2026

Kimi K3: the largest open model ever released

Moonshot AI released Kimi K3 — 2.8 trillion parameters, 1M-token context, native vision — beating the proprietary flagships of the previous generation on coding and agentic benchmarks at launch. For eleven days access ran only through Moonshot’s own servers — not an option for regulated data.

July 27, 2026

The weights land

Moonshot published the 1.56 TB Kimi K3 checkpoint. A frontier-tier model became a static file you can checksum, air-gap, fine-tune, and own — with no US or Chinese provider anywhere in the chain. The organizations that benefited on day one were the ones whose infrastructure was already racked.

August 25 – 31, 2026

Five frontier-class open-weight releases in nine days

GLM-5.3-Flash (MIT, multimodal, single-node), Qwen3.8-Flash-Next, GLM-5.3, Tencent Hy4-preview (Apache 2.0) and DeepSeek V4-Flash-Vision (MIT) all shipped weights in one week. The capability premium for rented APIs collapsed; licensing became the differentiator. Getting regulated organizations ready for exactly this moment is the job this practice was built to do.

Operating principles

Three rules we don’t bend.

Compliance by architecture

Residency clauses in a vendor contract are promises; data that physically never leaves your infrastructure is proof. Every deployment is designed to support our customers’ obligations — PIPEDA, Québec Law 25, PHIPA, OSFI Guideline E-23, law society confidentiality rulesHIPAA, GLBA, CMMC/ITAR, and state privacy acts— with zero data egress as the default posture, not an add-on.

Honest benchmarks, honest sizing

Full-precision Kimi K3 is supercomputer-class, and we say so — most organizations are better served by a single-rack GLM-5.2 deployment. We publish real hardware requirements, quantization trade-offs, and cost curves, and the sovereignty assessment tells you which tier your workloads actually justify — whether or not you build with us.

Model-agnostic, always

We have no model to sell you. Kimi K3, GLM-5.2, DeepSeek, Qwen, Llama, Mistral — the bench changes monthly, and your infrastructure is built to outlast every entry on it. Where procurement policy requires Western-origin weights, we deploy them. Sovereignty means model origin is a governance choice you get to make.

Practice areas

One firm, end to end.

Sovereign AI fails when it’s treated as an IT project. We run it as a complete practice — four disciplines under one engagement.

Build

GPU cluster design and procurement sized to your models and concurrency — on premises, in-country colocation, or fully air-gapped.

On-premise deployment →

Fine-tune

Parameter-efficient tuning on your corpus, trained inside your environment and evaluated against a harness built from your real work.

Fine-tuning & customization →

Host

Production in-country inference with SSO, role-based access, audit logging, and observability. Zero third-party API calls.

Hosting & maintenance →

Maintain

Open-weight releases arrive monthly. We evaluate them against your workloads, migrate fine-tunes, and upgrade serving stacks on your schedule.

Air-gapped AI →

Why we publish

Research transparency is the sales pitch.

Everything we learn deploying open-weight models — real hardware requirements, quantization trade-offs, compliance architectures, cost breakevens — gets written up and published in our guides and deployment blueprints. Regulated buyers deserve to check our reasoning before they trust us with their infrastructure, and honest technical writing is how a young practice earns that trust. If our numbers are wrong, tell us — we’ll correct them and date the correction.

The practice

Who does the work.

Chels.ai’s practice combines AI-infrastructure engineering, model evaluation and fine-tuning, and regulated-industry compliance architecture under one roof — operating from Toronto, serving clients Canada-wideacross the United States. Engagements are led end-to-end by the people who design the architecture, not handed off to a rotating bench.

Team profiles are being added as the practice formalizes — for now, the fastest way to evaluate us is to read the published work or get us on a call.

Start with the sovereignty assessment.

Two weeks. We map your data obligations and workloads, size the hardware, and hand you a written architecture with a real cost model — whether or not you build with us.

LATEST — GLM-5.3-Flash (MIT) · GLM-5.3 · Qwen3.8-Flash-Next · Aug 2026OUTPUT — written architecture + cost model vs. cloud spendTIMELINE — 2 weeks · fixed feeJURISDICTIONCanada-wide, from TorontoUS-wide deployments