Alibaba's open-weight Qwen 3.8 family is getting a serious speed upgrade. Qwen 3.8 27B, the dense multimodal model built for agentic coding, tool use, research and long-running workflows, is now available on Cerebras public inference endpoints. It accepts text and image inputs, supports configurable reasoning, and the weights are up on Hugging Face for teams that prefer to self-host.

The headline number is speed. Cerebras lists the model at around 1500 output tokens per second, far beyond what typical local GPU serving delivers for a model this size. Context runs to 64k tokens on the free tier and 128k on paid tiers, with up to 40k tokens of output, so long agentic tasks stream back quickly instead of stalling. Cerebras wafer-scale hardware does the lifting, turning a strong open-weight model into a genuinely low-latency API.

For founders, this is an agent economics story. Fast output tokens mean multi-step loops that verify their own work, browse sources and call tools feel responsive rather than sluggish, while open-weight pricing keeps per-task cost low. The launch raced to the top of Hacker News with hundreds of comments within a day, a clear signal of developer appetite for fast open inference. As Cerebras keeps adding models at this throughput, expect the convenience gap between open weights and proprietary APIs to keep narrowing.