← See all
May 2026 - Jun. 2026Open Source

mini-vla

What's the smallest VLA model viable for which task?

An ongoing research project: how small can a vision-language-action (VLA) model be while still learning to follow a language command and act on it, and how far does that hold across different tasks?

Done so far. A tiny language-conditioned policy learns, by behavioral cloning against an analytical inverse-kinematics expert, to pick up a named block on a 2-link arm ("grab the blue cube"). It's small enough to train from scratch live in the browser in under 60 seconds. You can try it at the top of this page.

Next. Moving past this hand-built pickup task: training a behavioral-cloning VLA policy in ManiSkill, a GPU-accelerated robot simulation benchmark, to see how well the same small-model approach transfers to more realistic manipulation tasks.

VLM