Zhipu's GLM line built its reputation on function calling and structured output rather than benchmark scores. Why that reliability is harder to achieve than raw capability, and where it matters.
OpenAI's GPT-5.6 ships as three tiers with different capability and cost. Which one you actually want for professional writing, and why the answer is usually not the flagship.
OpenAI ships gpt-oss under Apache 2.0 in 120B and 20B sizes. Where they fit, why a reasoning-oriented model is a mixed blessing for rewriting, and how they compare to the small Qwen and Gemma models.
A decision procedure that starts with your actual RAM instead of a leaderboard. Which size, which variant, which quantization, and how to know when you have picked wrong.
Four frontier models in two months and a new open-weights release most weeks. A filtering system for staying current on AI without turning it into a second job.
Model cards and system cards are the closest thing to a datasheet AI has. What to look for, what the omissions tell you, and why this is becoming a compliance document.
Thinking Machines Lab shipped its first open-weights model in July 2026. Why a US entry matters in a field that had tilted heavily toward Chinese labs.
Moonshot AI released Kimi K3 as a 2.8-trillion-parameter open-weights model in July 2026. What that actually means, who can run it, and why it changes the field even if you never touch it.
MiniMax M3 handles a million-token context at a fraction of standard transformer compute. How sparse attention works in plain terms, and why long context is cheaper but not better.
For European organizations, where a model runs is a compliance question. How Mistral's open lineup fits, and why sovereignty is solved by architecture more often than by geography.
NVIDIA keeps shipping compressed open-weight models derived from larger ones. How pruning and distillation work, and why compressed models are the ones that reach your laptop.