/// UNIFIED MULTIMODAL MODEL · EDGE-FIRST

The omni model small enough to be yours.

One small model that understands and generates — built to fit in the memory of a base-spec laptop. No cloud. No queue. No meter running.

LIVE / TALK TO MOCHE IN THE PLAYGROUND [ ~8B CLASS ] [ SHIPS Q4 ] [ APPLE SILICON ]
8GB
minimum footprint
4
modalities generated
+0.2GB
cost of a native voice
Q4
quantization ship gate
/// MANIFESTO[01]

The frontier is racing to build giants in datacenters.
moche runs the other way.

Every year models grow, and every year they drift further from the machines people actually own. moche is a bet on the opposite corner of the map: one small unified model — understanding and generation in a single interface — that fits where you already work.

Not a demo of a distant future. A thing you download, and it's yours.

/// ANATOMY[02]

One brain. Four organs.
Stitched with learned connectors.

A frozen multimodal backbone does the thinking. Generation is grafted on through latent connectors — the brain's inner state flows directly into each decoder, no lossy text bottleneck in between. Speech lives inside the brain itself, as codec tokens.

[01]BrainReads, watches, listens. Decides what to make and holds the whole conversation in one context.4B MULTIMODAL / FROZEN
[02]VoiceSpeech generated by the brain itself, token by token. No separate TTS model to load.CODEC TOKENS / +0.2GB
[03]Hand1024px images through a learned connector — the brain's intent, not a paraphrased prompt.COMPACT DiT / ~1.6B
[04]MotionShort clips on hardware where video generation simply hasn't existed before.VIDEO DiT / ~1.3B
┌────────────────  moche · ~8B class · ships Q4  ────────────────┐
│                                                                  │
│   [ brain — 4B multimodal LLM, frozen ]                          │
│      │                                                           │
│      ├── speech codec tokens ─→ codec decoder     voice        │
│      ├── <gen_image> → queries → connector → DiT   hand         │
│      └── <gen_video> → connector → video DiT      motion       │
│                                                                  │
└──────────────────────────────────────────────────────────────────┘
/// FIT[03]

Runs where you are.

base spec
8GB
  • full understanding
  • native speech
  • image generation
  • video — not this tier
~4GB resident / MacBook Air ok
sweet spot
16GB
  • everything in 8GB
  • short video, sequential load
  • the tier video never fit before
~7GB peak / this is the point
headroom
24GB+
  • all organs resident
  • upgraded decoders
  • longer, larger, faster
~10GB+ / desks & studios

Distributed as MLX bundles for Apple Silicon. CUDA builds follow.

/// CRAFT[04]

Small isn't a port.
It's the design constraint.

Every stage of moche has one gate: survive Q4 quantization or don't ship. And the training data isn't scraped — it's farmed. An orchestrated generation pipeline produces candidates, vision-language critics score every one, and only the winners enter the training set.

Survival of the fittest —
as a data curation strategy.

THE DARWIN FLYWHEEL / MOCHE TRAINING PIPELINE
/// ROADMAP[05]

Four stages. Each one ships.

[STAGE 1]

A voice — teaching the brain to speak

Codec-token vocabulary expansion. Speech becomes something the model says, not something a plugin renders.

IN PROGRESS
[STAGE 2]

A hand — drawing from intent

Learnable queries and a connector wire the frozen brain to a compact diffusion decoder.

NEXT
[STAGE 3]

Motion — the missing capability

Short video generation on 16GB machines, where it has never fit.

PLANNED
[STAGE 4]

Your Mac — the whole point

Quantized MLX packages, one inference wrapper, tiered bundles. Download, own, play.

PLANNED