A new open source project called mini-AGI is climbing Hacker News today with a claim that cuts against how almost everyone uses language models. It is a byte level model that trains from scratch on a single GPU with just 8 GB of VRAM, and it never stops learning. Rather than holding every parameter in memory, mini-AGI keeps its weights as ordinary files on disk and pages them onto the card only when they are needed. That means the size of the model is bounded by free disk space, not by the memory on the graphics card. The model also assembles its own architecture as it goes: it grows new capacity when it runs short and prunes weights that nothing asks for.
The stated motivation is that every model you can actually own today was trained by somebody else and then frozen. You can fine tune around the edges, but you cannot keep training it on what you do day to day, because the moment you try, it forgets what it already knew. mini-AGI takes a different route. It reads a stream of characters one chunk at a time, takes a gradient step on each chunk, and serves requests through exactly the same code path it trains on, so there is no separate inference implementation to keep in sync.
The author is upfront that this is a toy scale experiment, not a frontier model, and the trained weights are not published yet. The first pass over the training corpus is still running and the weights will only be released once that pass completes, which the README estimates at a few weeks away. The interesting part for founders is the shape of the idea rather than the benchmark numbers: a personal model that lives on your own disk, keeps improving from data you feed it, and does not depend on an API bill or someone else's roadmap. If continual learning without catastrophic forgetting can work at toy scale on a laptop, it is worth watching whether the same trick survives a larger budget.
