Poolside releases Laguna S 2.1, a 118B open-weight model that tops SWE-Bench Multilingual and outperforms rivals several times its effective size
Poolside released Laguna S 2.1 on July 21, 2026, a mixture-of-experts coding model that climbed to the top of the SWE-Bench Multilingual public leaderboard at 78.5% and beat larger open-weight rivals while needing only a single Nvidia DGX Spark to run. Weights landed on Hugging Face the same day under the OpenMDW 1.1 license, which permits commercial use at no cost.
What happened
The model carries 118 billion total parameters but activates only 8 billion per token, keeping inference costs low relative to models with comparable benchmark scores. It supports a 1-million-token context window and ships in both thinking and no-thinking variants, per Poolside's GlobeNewswire press release.
On Terminal-Bench 2.1, the model scored 70.2% in thinking mode. On SWE-Bench Multilingual it reached 78.5%, topping the public leaderboard at launch. Both results placed it ahead of DeepSeek-V4-Pro-Max and Nvidia Nemotron Ultra, which have much larger active parameter footprints, per The Decoder. GPT-5.6 Sol, Claude Fable 5, and Kimi K3 remain above it on Terminal-Bench 2.1.
Poolside started training on May 22, 2026 using 4,096 Nvidia H200 GPUs. From that date to release took under nine weeks. The company used FP8 precision for reinforcement learning, a first for its lab, and covered 409,000 post-training environments spanning terminal tasks and software engineering workflows, per The Decoder. As a secondary result, the model independently solved Erdos Problem #397, a combinatorics problem previously solved only by the largest frontier models.
Beyond Hugging Face, the weights run through the Poolside API, OpenRouter, Baseten, and Vercel AI Gateway. Laguna S 2.1 is the company's third release in three months, following M.1 and XS.2 in April 2026.
Why it matters
The 8B-active-parameter architecture lowers the hardware bar for frontier-grade coding assistance in a meaningful way. Matching or beating DeepSeek-V4-Pro-Max on SWE-Bench Multilingual previously required multi-node GPU clusters; Laguna S 2.1 does not, per The Decoder. A single DGX Spark sits within reach of most enterprise AI teams, which shifts the build-vs-buy math for shops running dedicated coding agents on-premises or in a private cloud.
The SWE-Bench Multilingual result carries more signal for polyglot codebases than Python-only evals. Teams working across TypeScript, Go, Java, or Rust get a benchmark that reflects their actual workload mix rather than a Python-dominated test suite.
Three models in three months also signals something about Poolside's cadence. Each release has posted higher benchmark numbers than its predecessor, and the lab now holds a public leaderboard position above both a Chinese open-weight frontrunner and Nvidia's own coding model, per Poolside's press release.
What to watch next
Third-party replication from Artificial Analysis and LiveBench will confirm whether the SWE-Bench Multilingual and Terminal-Bench 2.1 scores hold under independent testing conditions. Poolside has not yet published enterprise API pricing for Laguna S 2.1, so total cost of ownership for production deployments beyond self-hosting remains unclear.
Sources
- Poolside releases Laguna S 2.1, the West's most capable open-weight model: Poolside / GlobeNewswire, July 21, 2026
- Poolside's Laguna S 2.1 is a small open-weight coding model that punches well above its size: The Decoder, July 23, 2026
