Hardware-efficient building blocks for modern LLMs with linear attention, state space models, and hybrid architectures