AMD snaps up Taalas, the startup etching AI weights into silicon
AMD snaps up Toronto startup Taalas, whose model-specific chips etch weights directly into silicon and reportedly hit nearly 17,000 tokens per second in early demos.
BY FOUNDERBUILT AI NEWS
AMD has acquired Taalas, the Toronto startup that etches AI model weights directly into silicon. Announced at market close on Thursday, the deal is an actual acquisition rather than an acquihire, though terms were not disclosed. It lands in similar territory to Nvidia multibillion-dollar licensing deal with Groq: making premium inference services for AI agents, such as code assistants, faster and cheaper to run.
Taalas chips are model-specific integrated circuits that skip HBM entirely, storing weights in a mask-ROM recall fabric and KV caches in SRAM. Its first test chip, fabbed on TSMC 6nm process, served Meta Llama 3.1 8B at nearly 17,000 tokens per second, claimed to be 48x faster than Nvidia GPUs at the time. The second-gen HC2 chip, due this summer, targets 20 billion parameters per accelerator, meaning 50 chips could support a trillion-parameter model.
AMD intends to pair the technology with its Instinct-based Helios racks, handling prompt processing on GPUs while offloading token generation to Taalas accelerators. The catch: chips are locked to the models they are etched for, so large model changes require a re-spin, though only two layers of metal need altering. Still, etching weights into silicon is reportedly 100x cheaper than training a frontier model.