AI21 logra una reducción del 83 % en el tiempo de inicio de las cargas de trabajo de IA con AI Hypercomputer
Story summary
Nota del editor: AI21 Labs es un laboratorio de IA líder a nivel mundial con una larga trayectoria en la creación de modelos básicos, en particular la familia Jamba, y hoy se centra en LLM especializados y tecnología de optimización de agentes. Al adoptar Google Cloud AI Hypercomputer, AI21 redujo los tiempos de espera de trabajos de alta prioridad de 72
📌 Key Highlights & Takeaways
- Nota del editor: AI21 Labs es un laboratorio de IA líder a nivel mundial con una larga trayectoria en la creación de modelos básicos, en particular la familia Jamba, y hoy se centra en LLM especializados y tecnología de optimización de agentes.
- Al adoptar Google Cloud AI Hypercomputer, AI21 redujo los tiempos de espera de trabajos de alta prioridad de 72
Editor’s note : AI21 Labs is a leading global AI lab with a long track record of building foundation models, most notably the Jamba family, and today focuses on specialized LLMs and agent optimization technology. By adopting Google Cloud AI Hypercomputer, AI21 cut high-priority job wait times from 72 hours to 12 and manual scheduling interventions from 20 per week to zero.
At AI21 , we build foundation models and agent optimization products that help enterprises run agents at frontier quality, efficiently. Our language models, including the Jamba family, and our agent optimization product suite run demanding production workloads, including our own. We chose Google Cloud AI Hypercomputer to support them at scale.
To keep our model training runs highly utilized, we needed a performant, scalable environment codesigned across infrastructure, orchestration, and consumption models. Our model training runs on one of our shared Google Kubernetes Engine (GKE) clusters, pooling thousands of Google Cloud A3 (powered by NVIDIA H100 Tensor Core GPUs) and A3 Ultra (powered by NVIDIA H200 Tensor Core GPUs) instances, so any team can draw on the full capacity of the fleet rather than being boxed into its own slice. The cluster also trains models and agent-optimization workloads beyond the Jamba family. That approach keeps utilization high, and it makes scheduling hard.
Prior to leveraging GKE for orchestration, we used to negotiate capacity by hand in Slack. If you needed capacity for a training run, you posted in #gpu-resources and hoped for the best.
That worked fine when the cluster had headroom. It stopped working once utilization pinned near 100%, which is where you want a reserved compute fleet to sit.
Over time, every request became a negotiation. Team leads spent their time refereeing compute disputes. Our high-priority jobs — the large, multi-node training runs that need half or more of the cluster at once and serve as the critical path for model projects — could sit blocked for up to 72 hours waiting for enough contiguous capacity to open up.
Scarcity created two distinct problems, and it took us a while to see them as separate. The first was contention: determining who gets compute access next, which we resolved through negotiation. The second was fragmentation: capacity that was technically free but scattered in pieces too small for a large job to use, a bin-packing problem no amount of negotiation could fix.
Sometimes we had plenty of capacity free on paper, but it was scattered across different machines in chunks too small for a larger job to actually land. Without all-or-nothing admission, the cluster could reach a deadlock, with machines holding resources without doing useful work until someone stepped in manually.
Cryptographic Security & Key Generator
Generate entropy-tested high-security keys and encryption-grade tokens.
Source: Cloud Blog.
Read the full story at the original source ↗
For questions: mrsmithcons@gmail.com.
☁️ Complete Cloud Credit Application Guide & Architecture Specs
Direct application templates, fast-track partner codes, and architecture benchmarks.
⚡ Access Cloud Playbook ➔