Startup Taalas has unveiled hard-wired silicon that runs Llama 3.1 8B at 17,000 tokens per second while cutting costs by 20x. Can custom chips beat GPUs?
Models & Hardware DeskDRAM chip prices have jumped 7x over the past year, forcing AI developers to optimize memory orchestration and prompt caching to control inference costs.
Models & Hardware Desk