Companies

AMD Buys Taalas, the Startup Baking Models Into Silicon

AMD has acquired Taalas, which etches AI models directly into silicon. The demo chip hits 16,000 tokens per second, at the cost of locking each chip to a single model.

Muhammet Fatih BatmanAugust 8, 20263 min read5 views
AMD Buys Taalas, the Startup Baking Models Into Silicon

Every AI hardware announcement arrives with a speed record attached. What stands out in this week's acquisition is not the speed itself but the route taken to get there.

On August 7, AMD announced it had acquired Taalas, a Toronto company whose entire business fits in one sentence: it burns a model's architecture and trained parameters directly into the chip. Terms were not disclosed, and the deal is subject to standard regulatory approvals.

What exactly did AMD acquire?

Running a model today is a two-part arrangement: a general-purpose chip, usually a GPU, plus the model files you load onto it. Taalas collapses that split by etching the model into silicon at manufacturing time. The result is a chip that runs one model extraordinarily fast and cannot run any other model at all.

The company was founded in Toronto in 2023 and came out of stealth in February 2026. Co-founder Ljubisa Bajic is a familiar name in silicon design circles. Its demo chip produced more than 16,000 tokens per second per user on Llama 3.1-8B, many times what the same model manages on standard accelerators.

Vamsi Boppana, AMD's senior vice president for AI, said the technology will be folded into the company's accelerator roadmap and offered alongside Instinct GPUs as a system-level solution. No product timeline was given.

The trade-off is not hidden. A model baked into hardware cannot be updated. When a new version of the model ships, the hardware has to ship with it. Google is reportedly working along similar lines under the name "Frozen v2," so AMD is not alone in taking this bet.

Why this matters if you never buy a chip

The short answer is your invoice. Most companies do not build their own inference hardware. They buy AI through an API, priced per token. And what sets that token price is not training, it is inference: the compute burned every time the model answers a question.

In our own client work, the AI projects that get shelved rarely fail because the model was not good enough. They fail because volume grew and the monthly bill stopped being predictable. Any hardware move that pushes inference cost down also lowers the threshold at which the business case starts to work.

Put rough numbers on it. A support assistant handling 5,000 customer questions a day at an average of 1,500 tokens per question burns roughly 225 million tokens a month. At that volume, a 30 percent drop in token price is a line item you can see in the budget. It is also, in our experience, the difference between a project that graduates from pilot and one that quietly stays there.

That is where Taalas makes sense, and where it does not. One fixed, heavily used model is the ideal case: a support assistant, a document classifier, a product search engine. A team that swaps models every quarter would be trading away exactly the flexibility it depends on.

None of this shows up on a 2026 invoice. It shows up in 2027 and 2028 budgets. But the direction is legible: the industry is shifting from a race to run the biggest model toward a race to run the same job for less.

Sources: The Decoder

Share This Article

Muhammet Fatih Batman

Written by

Muhammet Fatih Batman

Founder & Editor

Founder of YZ Uzman, with 20+ years of experience in web design and software development.

More news

Want to put this technology to work in your business?

Let's talk