The frontier is racing to build giants in datacenters.
moche runs the other way.
Every year models grow, and every year they drift further from the machines people actually own. moche is a bet on the opposite corner of the map: one small unified model — understanding and generation in a single interface — that fits where you already work.
Not a demo of a distant future. A thing you download, and it's yours.
One brain. Four organs.
Stitched with learned connectors.
A frozen multimodal backbone does the thinking. Generation is grafted on through latent connectors — the brain's inner state flows directly into each decoder, no lossy text bottleneck in between. Speech lives inside the brain itself, as codec tokens.
┌──────────────── moche · ~8B class · ships Q4 ────────────────┐ │ │ │ [ brain — 4B multimodal LLM, frozen ] │ │ │ │ │ ├── speech codec tokens ─→ codec decoder voice │ │ ├── <gen_image> → queries → connector → DiT hand │ │ └── <gen_video> → connector → video DiT motion │ │ │ └──────────────────────────────────────────────────────────────────┘
Runs where you are.
- full understanding
- native speech
- image generation
- video — not this tier
- everything in 8GB
- short video, sequential load
- the tier video never fit before
- all organs resident
- upgraded decoders
- longer, larger, faster
Distributed as MLX bundles for Apple Silicon. CUDA builds follow.
Small isn't a port.
It's the design constraint.
Every stage of moche has one gate: survive Q4 quantization or don't ship. And the training data isn't scraped — it's farmed. An orchestrated generation pipeline produces candidates, vision-language critics score every one, and only the winners enter the training set.
Survival of the fittest —
as a data curation strategy.
Four stages. Each one ships.
A voice — teaching the brain to speak
Codec-token vocabulary expansion. Speech becomes something the model says, not something a plugin renders.
A hand — drawing from intent
Learnable queries and a connector wire the frozen brain to a compact diffusion decoder.
Motion — the missing capability
Short video generation on 16GB machines, where it has never fit.
Your Mac — the whole point
Quantized MLX packages, one inference wrapper, tiered bundles. Download, own, play.