StrataFlow
Run very large models on the hardware you already own - no GPU required. StrataFlow is a local inference engine that streams Mixture-of-Experts weights from disk, so a model far bigger than your RAM runs on a normal laptop or desktop CPU inside a bounded memory budget.
What it is
StrataFlow is a local inference engine for large Mixture-of-Experts (MoE) models. A MoE model is huge on disk but only uses a small slice of its weights for any single token. StrataFlow keeps the always-on trunk plus a bounded number of active experts resident, and streams the rest in from disk per token.
It runs its own forward pass over ggml. StrataFlow is not a wrapper that hands your model to another engine - it builds and executes the compute graph itself, so it decides exactly which weights are in memory and when.
How a model bigger than my RAM runs
The question everyone asks: the model is 24 GB, how can it run on 12 GB (or 8 GB) of RAM? The answer is that RAM holds only the working set, not the whole model. Weights live in strata - layers of memory - and flow between them on demand:
VRAM (fastest, smallest) <- hot experts, trunk, KV cache (optional GPU)
^v
RAM (fast, medium) <- warm experts, resident trunk
^v
SSD (large, slowest) <- the full model; cold experts stream in per tokenFor each token, only the trunk and the top-k active experts per layer are needed. StrataFlow keeps just that working set in RAM and streams the cold experts from disk. Everything else only has to be reachable on disk, not held in memory. You set the resident budget with a single knob (--expert-slots), and resident RAM stays bounded even as the on-disk model grows.
More memory buys speed, not capability
The output is identical at every memory size. A bigger RAM budget just lets more experts stay cached, so fewer disk reads are needed per token. The answer the model produces does not change.
What actually constrains an 8 GB machine
The real limits are (1) enough disk space for the model file and (2) streaming speed - your SSD read throughput sets tokens per second. RAM size does not cap what you can run.
Why it is different
Its own forward pass over ggml
StrataFlow builds and runs the compute graph itself, so it controls residency directly. It is not a skin over another engine.
One tiered weight store
A single store across VRAM, RAM and SSD with a shared caching policy - not three separate mechanisms bolted together.
True per-expert streaming
Each layer's router is read first, then only the top-k active experts are made resident before the expert matmul runs - so resident expert memory is bounded to what a token actually uses, not the whole layer.
Automatic hardware tuning
It profiles your machine on first run and decides what goes where. No hand-tuning placement flags.
A streaming-friendly format (.strata)
Built from standard GGUF, one file holds the always-resident trunk plus aligned per-expert blobs for fast per-expert reads.
Predictor-driven prefetch
The router read-back feeds a predictor that warms the next layer's likely experts into the cache ahead of use, raising the cache hit rate without changing results.
Proof points (measured)
StrataFlow's correctness and bounded-RAM behavior are measured, not asserted.
Byte-identical to the reference engine
Every engine change is gated on reproducing the reference logits bit-for-bit (within floating-point noise), so StrataFlow's own forward pass matches a standard decode.
Bounded RAM, flat as the model grows
A 785 MiB model runs at ~132 MiB peak RSS, and a 1177 MiB model also runs in ~132 MiB - peak memory stays flat as the on-disk model grows.
Quantized weights validated
Q8_0 and the K-quant families Q4_K/Q6_K are validated against the oracle (identical greedy token sequence) and stream through the .strata path.
A benchmark harness reports real numbers
Tokens per second, time to first token, peak RSS, resident weights and bytes per token - across a matrix of model sizes and memory budgets.
The bounded-RAM mechanism is proven, and a real downloaded quantized model runs end to end through it. Numbers for a real large MoE are being gathered on real hardware - see Current status below.
Who it is for
Current status
StrataFlow has a working engine core on CPU, green in CI on Linux, macOS and Windows. We keep this honest - the mechanism is proven, and real large-model numbers are being gathered.
Working today
Honest open items
Platforms
Windows, Linux and macOS
The engine core is green in CI across all three platforms.
CPU today, GPU optional
CPU is the no-GPU mission. A GPU, if present, is an optional accelerator for speed - never a requirement. GPU backends are coming.
Try it
StrataFlow runs on CPU. The easiest way to see the big-model, bounded-RAM, no-GPU path on a real machine is the Colab guide, which builds the engine, packs a model to .strata, and streams it with a bounded memory budget.
guide: colab/README.md in the repository
Licensing and commercial use
StrataFlow is source-available, under the Coaade Source-Available License, Version 1.0.
The owner and licensor is Coaade Inc., a Delaware C corporation. StrataFlow is built on llama.cpp / ggml (MIT), and the repository contains no model weights. Want to use StrataFlow commercially, or discuss terms? Contact contact@coaade.com.
Run large models locally
Big model, bounded RAM, no GPU. Reach out for commercial licensing, or explore the rest of the Coaade product line.