Free Transformer: Can Latent Variables Improve LLM Reasoning?
The Free Transformer adds learned latent variables to decoder LLMs. Here’s how it works, what benchmarks show, and why builders should care.
Jul 25, 202620 min read
3 articles on transformers.
The Free Transformer adds learned latent variables to decoder LLMs. Here’s how it works, what benchmarks show, and why builders should care.
Attention residuals turn Transformer depth into a retrieval system. See how Kimi’s Block AttnRes improves efficiency, reasoning, and what remains unproven.
Looped transformers reuse layers to reason in latent space. Here’s why they matter, where they outperform CoT, and what still limits them.