How it fits together
Six questions investors ask
Why does this need to be one company?
Because the pieces compound. Compression makes a 27B fit, rMLX makes it run, the Harness makes it reliable and useful, and E-RLM is the next step in how it thinks. Each attacks a different bottleneck.
Why wouldn’t the big labs just do this?
We don’t compete on weights. We build the layer that takes a model from any provider — DeepSeek, Kimi, Gemma, Qwen — and compresses, recovers, runs and deploys it on hardware people already own, with the product on top.
Where is the business?
Ready to start a closed pilot of the platform with businesses. LocallyAI AU, our subsidiary, installs and supports it: the machine, the on-site install, configuration and support.
What is the primary business?
Local install and configuration of on-device AI for Australian businesses, through trained contractors. Alongside it, the planned LocallyAI Platform: a private machine in our cluster, rented by the month.
Why so many industries?
Two reasons. We put learning first and apply it across the whole stack, so it stays one cohesive product. And the system is built to be adaptable and reliable, so it is genuinely useful across many industries.
Does the compressed model keep its intelligence?
Divergence 0.3737 inside the 0.40 gate, repetition 0.0394 inside 0.25, over 32,736 scored positions. A task-level benchmark against the full-precision model is the next receipt and the release gate.
The business
Studio C / pre-seed, September 2026
Frontier-capable AI, on ordinary hardware.
One stack, three bottlenecks, one product: capable models moved onto hardware you own.
Three co-founders: Harry, Matt and Riley. Studio C owns the technology; LocallyAI AU installs and supports it.
1 / 04
StudioC Harness / local-first workspace
Harness
Private AI that finds, understands and gets work done — on a machine in your building. Nothing leaves unless you switch it on.
300 / 300
structured tool calls with the Harness’s automatic correction, from 287 / 300 as written, on the stock 4-bit model.
The product / live footage of the interface, no concept renders
Valid structured tool calls, 300 prompts, the stock 4-bit model. The Harness recognises and corrects malformed calls automatically from a code library of malfunctions captured over thousands of tool calls, part of the Forge compress-and-recover pipeline. No weight changed. Also 12 of 12 on an Exercism-based code set.
2 / 04
Project Anneal / complete package size, GiB, same Qwen3.8-27B
Same model.
Smaller by 88.5%.
Measured on the 7.58 GiB runtime
Project rMLX / against MLX 0.32.2
A full Rust rewrite of Apple’s MLX. Parity 174 / 174 first. Paired medians, n=14, one M1 Max, $0 of cloud. Two gates are open: the fused chain misses our own 0.47 µs bar, and decode is 408.5 tok/s against MLX’s 437. Released as open code under the same terms as the Harness.
The 5.769 GiB build is untrained; the 7.581 GiB runtime carries the quality numbers. Ridge’s figure is third-party, from its model card. Every size is the whole package.
Pre-seed / September 2026
27B.
5.77 GiB.
8.68× smaller than the original weights. Complete, 1.5-bit ternary.
The only runtime for this model that stays usable on the machine, with room left for the OS and an agentic context window. Retains an estimated 80% of knowledge; we are working to 90% or more before it ships.
A hybrid rounding-encoding method, different to standard quantisation, with a short recovery run (3–24 hours by model size) in our Forge pipeline. It applies to any model. The Harness recovers what compression costs, correcting known malformations at kernel level: this compressed 27B goes from 0 of 15 structured tool calls as written to 15 of 15 with automatic correction. It is not shipped yet; the Harness runs stock models until then. Forge, the engine behind it, has also found a 0.262-bit compression on the MLP blocks it has reached so far and is iterating on it actively; Forge itself is not planned for release.
Coming next
1B · 2 bits4B · 2 bits8B · 1.5 bits14B · 1.5 bits35B-MoE · 1.5 bits80B-MoE · 1.5 bits180B-MoE · 1.5 bits
Also planned: DeepSeek V4 Flash, Kimi K3 and other frontier models compressed for non-commercial use, with runtimes for enthusiast and business hardware; Gemma 4 (26B and 31B) and Qwen3.8-27B, Qwen3.8 180B and Qwen3-Next 80B-A3B on everyday consumer hardware and phones; 1B to 14B and other sizes on edge devices.
Neither the encoding nor the recovery pass needs the original training data. Runtime today: our own; LM Studio and llama.cpp cannot run it yet.
3 / 04
Original architecture
E-RLM
Executable Reasoning Language Model.
The runtime answers the model’s arithmetic and runs its code while it is still thinking. No request, no wait, no reply.
Early research, not in any product yet. Capability numbers stay internal for now.
Two live examples / a person types at 8 tokens a second, the model answers at 15, the runtime fills in orange
4 / 04